A new atlas maps 2,226 coding tasks across 11 AI benchmarks

zainhas · x · 2026-07-24

A benchmark atlas maps 2,226 coding tasks across 11 eval suites

This post points to a new empirical atlas of AI coding benchmarks. The project embeds, maps, and classifies 2,226 unique tasks drawn from 11 benchmarks — the five used in Poolside's Laguna S 2.1 trajectory archive plus six community benchmarks.

The goal is to answer three questions:

The screenshot shows the dashboard behind the atlas, including benchmark clustering, task search, and per-benchmark breakdowns. It also highlights how the dataset separates problems by benchmark family and surfaces missing categories such as test-writing, code review, and Ctasks.

Related event: New Tool Maps Out 11 AI Coding Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →