Tencent’s WorkBuddy Bench evaluates coding agents across code, web, office and security
tencent · hf · 2026-07-24
Tencent releases WorkBuddy Bench for multi-domain coding agents
Tencent WorkBuddy Bench is a new evaluation suite for coding agents that covers four work domains: Code, Web, Office, and Security.
The benchmark is designed to reduce contamination by reverse-engineering each task from a real commit, pull request, or business scenario, then rewriting it as a short colloquial request that cannot be recovered by searching public issue threads. The release is fully open and includes:
- task directories
- environment images
- evaluation harnesses
- tests
- reference solutions
Because each subset uses a different scoring instrument, Tencent says scores are not comparable across subsets, so the suite does not report a single average. The paper also includes a cross-model leaderboard evaluated on two harnesses: CodeBuddy Code and Claude Code.
More from coding & agent
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11
- Dev builds browser 3D pizza delivery game with Claude: physics, GPS pathfinding, traffic AI — vinishkapoor · 2026-09-11
- Build X Carousel Posts from One Wide Image: A Splitter Tool Plus YouMind Skill Workflow — sujingshen · 2026-09-11
- "Anyone still coding the old way?" The joke capturing post-AI programming culture — lxfater · 2026-09-11