Switch language한국어
Back to the list

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

TL;DR AI

Key summary

2 min read
  1. Tencent released WorkBuddy Bench, an open benchmark for evaluating coding agents across code, web, office, and security tasks.

  2. The tasks are drawn from real commits, pull requests, and business scenarios, then rewritten to reduce prompt contamination.

  3. Tencent also released task files, environments, tests, and reference solutions to support reproducible scoring.

  4. The benchmark includes a cross-model leaderboard, making it easier to compare agents in realistic work settings.

Read the original