Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score

almmaasoglu · x · 2026-09-23

Developer Alm Maasoglu argues that current code benchmarks are mostly slop because they collapse very different workloads into a single score, hiding real capability differences.

He calls for narrower evals targeting specific domains: research code, backend APIs, database systems, infra, debugging, and repo-scale changes. "Good at coding" is not a useful capability category, he says, urging someone to fix the benchmarking status quo — echoing widespread community frustration with low-resolution SWE-style leaderboards.

Original post →

More from Models

Models channel →