Benchmarks should publish coverage boundaries: a coding agent fixing a bug isn't the whole job

Shahules786 · x · 2026-09-03

Drawing on CRMArena-Pro and "Designing Benchmarks for Knowledge Work", this thread argues leaderboards must state what they don't measure: an agent that fixes a bug hasn't handled requirements, review, docs, deployment or handoff. It proposes a four-field reporting standard — represented activity, tested setting, required work product, evaluated result. CRMArena itself found SOTA LLM agents complete under 40% of realistic CRM tasks with ReAct prompting.

Related event: Harvard/Stanford Paper Proposes Four-Field Reporting Standard for Knowledge Work Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →