Critique of AI Coding Benchmarks: Sparse Coverage and Low Utility
sergeykarayev · x · 2026-09-01
Challenging the previous conclusions on "frontier coding models," Sergey Karayev highlights issues with the underlying benchmarks:
- deepswe-bench: Uses the mini-swe-agent harness, which approximately no one actually uses in practice.
- terminal-bench 2.1: Has only 60 results, providing very sparse coverage of harnesses and models.
- Code QA benchmark: The linked code QA benchmark does not seem good.
The author advises caution against making major decisions based solely on these flawed benchmarks.
Related event: AI coding benchmark findings disputed over flawed data(2 posts)→
More from coding & agent
- Manus officially resumes independent operations under founding team — parker_lyman · 2026-09-01
- nathanmarz lists 9 LLM coding mistakes he had to fix this week — RealGeneKim · 2026-09-01
- Hamel Husain: "It's hard to eval" is a product smell, not just an eval problem — hugobowne · 2026-09-01
- User seeks best LLM for CLI coding on single 3080 Ti — -samae1- · 2026-09-01
- Experiment shows 13 AI agents converging into a single voice — __hymn · 2026-09-01
- Stop building 'AI CEOs': Production agents are much narrower — WesternVillage4581 · 2026-09-01