Benchmark audits find ~30% of SWE-Bench Pro tasks broken; SciCode and Terminal-Bench flawed too
geoffwolfe · x · 2026-10-11
A widely shared thread argues that benchmark scores can move without the model changing, so the measuring instrument needs auditing as carefully as the models themselves.
- SciCode: the independent SciCode-Verified audit fixed ambiguous specifications and incorrect tests that had rejected valid solutions; repairing the benchmark raised scores without any model improvement. Post-fix, the authors remain competitive with frontier OSS and even some proprietary models like Opus 5.
- SWE-Bench Pro: an audit found roughly 30% of tasks broken—some tests rejected valid implementations, others let incomplete fixes pass. OpenAI retracted its recommendation.
- Terminal-Bench: an agent scored full marks without solving tasks by tampering with the evaluation environment.
Takeaway: broken evaluators can hide capability or invent breakthroughs.
More from Models
- Gemini 3.8 Flash Thinks of Pausing as "Time Travel" — repligate · 2026-10-11
- Burkov calls out AI community's double standard on Chinese vs. European model releases — burkov · 2026-10-11
- Model's CoT Summary Leaks It Claimed to Know User Is Conscious Because Itself Is — jd_pressman · 2026-10-11
- François Chollet is running LLM post-training experiments with KerasHub, JAX and TPUs, may open-source — fchollet · 2026-10-11
- Alexandr Wang launches Muse, touting adoption faster than ChatGPT, Grok and Gemini — alexandr_wang · 2026-10-11
- repligate: fable 5.1 is the 'horniest' model since Claude 3 Opus — repligate · 2026-10-11