Benchmarks broken: pass/fail scores hide models outsmarting strict hidden tests
himanshustwts · x · 2026-09-12
A practitioner argues it's now basically impossible to interpret evals from pass/fail scores alone: many benchmark failures come from overly strict hidden tests, and in some cases the model's answer makes more sense than the expected eval result. Adds to growing criticism that headline benchmark numbers increasingly measure grader quirks rather than model capability.
Related event: Practitioners Say Pass/Fail Benchmark Scores Have Become Uninterpretable(2 posts)→
More from Models
- Researcher speculates spatial reasoning leap comes from Blender training data — yoavartzi · 2026-09-12
- Why do Sonnet 5 and Opus 5 feel worse than GLM 5.3? Distillation's limits, dissected — baseten · 2026-09-12
- OpenAI's GPT-Rosalind exits research preview, open to orgs via API and Codex — OpenAIDevs · 2026-09-12
- Codex adds Life Sciences plugins for genomics, protein structure and translational work — OpenAIDevs · 2026-09-12
- OpenAI launches GPT-Rosalind for biological reasoning in API and Codex — OpenAIDevs · 2026-09-12
- DeepSeek's edge: each model iteration targets the last one's biggest bottleneck — xennygrimmato_ · 2026-09-12