Benchmark saturation was predictable: OTS training data now shapes around popular evals

Shahules786 · x · 2026-09-07

Shahul argues recent benchmark saturation was predictable: once an eval is widely adopted, selling off-the-shelf training data shaped around it becomes easier than building a new benchmark across a new domain × capability axis. He cites Terminal-Bench and ARC-like data as examples and asks which benchmark falls next.

Original post →

More from Research

Research channel →