Tencent’s WorkBuddy Bench evaluates coding agents across code, web, office and security
tencent · hf · 2026-07-24
Tencent releases WorkBuddy Bench for multi-domain coding agents
Tencent WorkBuddy Bench is a new evaluation suite for coding agents that covers four work domains: Code, Web, Office, and Security.
The benchmark is designed to reduce contamination by reverse-engineering each task from a real commit, pull request, or business scenario, then rewriting it as a short colloquial request that cannot be recovered by searching public issue threads. The release is fully open and includes:
- task directories
- environment images
- evaluation harnesses
- tests
- reference solutions
Because each subset uses a different scoring instrument, Tencent says scores are not comparable across subsets, so the suite does not report a single average. The paper also includes a cross-model leaderboard evaluated on two harnesses: CodeBuddy Code and Claude Code.
More from coding & agent
- Webcmd says browser agents waste tokens on navigation, not the actual task — heyshrutimishra · 2026-07-24
- AREX introduces a recursively self-improving deep research agent with inner and outer loops — _reachsumit · 2026-07-24
- A local open-source agent skill flags the rest of your GitHub diff — Saboo_Shubham_ · 2026-07-24
- Stop exposing thousands of tools to agents; use a code sandbox instead — edgestone22 · 2026-07-24
- “Letting Claude read my codebase is basically open-sourcing it,” says developer — _Stocko_ · 2026-07-24
- Roboto Agents adds root-cause analysis that links robot logs to source code — Scobleizer · 2026-07-24