Benchmarks should publish coverage boundaries: a coding agent fixing a bug isn't the whole job
Shahules786 · x · 2026-09-03
Drawing on CRMArena-Pro and "Designing Benchmarks for Knowledge Work", this thread argues leaderboards must state what they don't measure: an agent that fixes a bug hasn't handled requirements, review, docs, deployment or handoff. It proposes a four-field reporting standard — represented activity, tested setting, required work product, evaluated result. CRMArena itself found SOTA LLM agents complete under 40% of realistic CRM tasks with ReAct prompting.
More from coding & agent
- Revera: Lean-verified POSIX regex engine with identical output in 6 languages — jedisct1 · 2026-09-03
- Fable 5.1 produced a result in ~10 minutes on the $100 plan — jasondeanlee · 2026-09-03
- Runway launches Dev MCP server for coding agents — tlakomy · 2026-09-03
- 7 GitHub Repos Turn One AI Agent Into an OS: Memory, Model Routing, Free Compute Stacks — garrytan · 2026-09-03
- The Zvi: agents rewire your reflexes — annoyances now get fixed by just asking Claude Code — TheZvi · 2026-09-03
- "A smarter model in a bad system just makes expensive mistakes faster" — iamKierraD · 2026-09-03