A new atlas maps 2,226 coding tasks across 11 AI benchmarks
zainhas · x · 2026-07-24
A benchmark atlas maps 2,226 coding tasks across 11 eval suites
This post points to a new empirical atlas of AI coding benchmarks. The project embeds, maps, and classifies 2,226 unique tasks drawn from 11 benchmarks — the five used in Poolside's Laguna S 2.1 trajectory archive plus six community benchmarks.
The goal is to answer three questions:
- what current coding benchmarks actually measure
- which languages and task types they cover
- where the gaps and redundancies are
The screenshot shows the dashboard behind the atlas, including benchmark clustering, task search, and per-benchmark breakdowns. It also highlights how the dataset separates problems by benchmark family and surfaces missing categories such as test-writing, code review, and Ctasks.
Related event: New Tool Maps Out 11 AI Coding Benchmarks(2 posts)→
More from coding & agent
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11