Mark Tenenholtz: Unsat Benchmarks Over 3 Months Old Are Just Badly Built

marktenenholtz · x · 2026-10-07

Mark Tenenholtz argues that most benchmarks older than 2-3 months that aren't saturated only have headroom because the tasks are badly written, the grading is poorly implemented, or both — not because models genuinely can't do better.

Original post →

More from Models

Models channel →