Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score
almmaasoglu · x · 2026-09-23
Developer Alm Maasoglu argues that current code benchmarks are mostly slop because they collapse very different workloads into a single score, hiding real capability differences.
He calls for narrower evals targeting specific domains: research code, backend APIs, database systems, infra, debugging, and repo-scale changes. "Good at coding" is not a useful capability category, he says, urging someone to fix the benchmarking status quo — echoing widespread community frustration with low-resolution SWE-style leaderboards.
More from Models
- Claude Opus 5.5 wows users with creative output: 'draw everything you want' — repligate · 2026-09-23
- Early tester: Claude Opus 5.5 so good we mistook it for a new model — EricBuess · 2026-09-23
- Opus 5.5 tops all three performance metrics on ArtificialAnalysis.ai — Wsz2020 · 2026-09-23
- Heavy Opus 5.5 All-Day Use Barely Dents Usage Limits, Users Report — altryne · 2026-09-23
- Show today's LLMs to experts 10 years ago and they'd call it AGI — JacksonKernion · 2026-09-23
- Tired of running out of credits, this Reddit user says cheap model swarms work surprisingly well — AnotherWallace · 2026-09-23