Evals are increasingly unreadable: overly strict hidden tests mislabel good answers
trq212 · x · 2026-09-12
The author argues pass/fail scores alone no longer meaningfully interpret model evals: many observed benchmark failures stem from overly strict hidden tests, and in some cases the model's answer makes more sense than the expected eval result — pointing to eval quality itself as a growing noise source.
More from Models
- Arcee open-sources NAC agent harness; 489 tool calls, 67 passing tests on DeepSeek V4.1 Flash — MaziyarPanahi · 2026-09-12
- Ant's Ling-3.0-flash-VL Scores 25 on AA Index With Just 5.5B Active Params — ArtificialAnlys · 2026-09-12
- Open-weight competition from China blocks AI labs from Uber-style price hikes, argues developer — firasd · 2026-09-12
- Frontier labs' privacy terms are 'insane': one thumbs-up can void your chat protections — niloofar_mire · 2026-09-12
- Redditor argues Astra hides chain of thought to prevent distillation, not to improve the model — bfkill · 2026-09-12
- Cursor ships CursorBench 4.0; speculation swirls that Gemini 3.8 post-training differs — ivan_bezdomny · 2026-09-12