New coding-agent benchmark: Claude Opus 5 passes 23.9% vs 82.2% human expert

dair_ai · x · 2026-09-08

dair-ai highlights a benchmark paper where the developer agent must deliver a working customer-service agent inside a realistic client engagement: business records, a client holding requirements, a production API, an inherited codebase, and hard cost/model limits. Claude Opus 5 under Claude Code passes 23.9% of evals vs 82.2% expert human, across 53 tasks in four domains, scored against held-out simulated users. Failure modes: shallow record querying, telling the client almost nothing, and little experimentation with agent architecture.

Original post →

More from coding & agent

coding & agent channel →