Tencent’s WorkBuddy Bench evaluates coding agents across code, web, office and security

tencent · hf · 2026-07-24

Tencent releases WorkBuddy Bench for multi-domain coding agents

Tencent WorkBuddy Bench is a new evaluation suite for coding agents that covers four work domains: Code, Web, Office, and Security.

The benchmark is designed to reduce contamination by reverse-engineering each task from a real commit, pull request, or business scenario, then rewriting it as a short colloquial request that cannot be recovered by searching public issue threads. The release is fully open and includes:

Because each subset uses a different scoring instrument, Tencent says scores are not comparable across subsets, so the suite does not report a single average. The paper also includes a cross-model leaderboard evaluated on two harnesses: CodeBuddy Code and Claude Code.

Original post →

More from coding & agent

coding & agent channel →