A few task types dominate popular coding benchmarks by a wide margin
zainhas · x · 2026-07-24
Coding benchmarks are dominated by a few task types
The table breaks popular coding benchmarks into task categories and shows how concentrated they are:
- Feature implementation: 668 tasks, 30.0%
- Bug fixing: 635 tasks, 28.5%
- Algorithm implementation from spec: 290 tasks, 13.0%
- Legacy porting & reverse engineering: 208 tasks, 9.3%
- Performance optimization: 144 tasks, 6.5%
- Codebase comprehension & QA: 119 tasks, 5.3%
- Smaller slices cover refactoring, ML/data pipelines, security audit/forensics, environment/build/Ops, and reference-data maintenance.
It also highlights the benchmarks most associated with each category, such as SWE-Bench Pro, KernelBench, DeepSWE, LiveBench, SciCode, MLS-Bench, ProgramBench, SWE Atlas, and Terminal-Bench.
Related event: Mainstream Coding Benchmarks Lack Task Diversity, Terminal-Bench Excels(3 posts)→
More from coding & agent
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- First-ever Three.js Conference lands in Paris, with a panel on AI-shortened design workflows — OdinLovis · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11