New coding-agent benchmark: Claude Opus 5 passes 23.9% vs 82.2% human expert
dair_ai · x · 2026-09-08
dair-ai highlights a benchmark paper where the developer agent must deliver a working customer-service agent inside a realistic client engagement: business records, a client holding requirements, a production API, an inherited codebase, and hard cost/model limits. Claude Opus 5 under Claude Code passes 23.9% of evals vs 82.2% expert human, across 53 tasks in four domains, scored against held-out simulated users. Failure modes: shallow record querying, telling the client almost nothing, and little experimentation with agent architecture.
More from coding & agent
- Crimsonland decomp hits 70.5% matched, crowd-sourced with AI coding agents — banteg · 2026-09-08
- Dev builds a silly 3D game, waits for codex quota reset to do it in an hour — jh3yy · 2026-09-08
- Matt Shumer publishes full GPT-6 Astra 3D workflow after builds top 15M views on X — mattshumer_ · 2026-09-08
- From one prompt to photoreal Manhattan: Matt Shumer's five-step agent loop for GPT-6 Astra 3D worlds — mattshumer_ · 2026-09-08
- Hospital multi-agent system suggested discharging a patient to free a bed — and the PHI visibility dilemma — ---starboy-- · 2026-09-08
- Truffle Journal: iPhone Markdown app that ChatGPT saves to, searches, and updates via MCP — Truffle_Journal · 2026-09-08