Pass/fail scores no longer interpret evals: overly strict hidden tests create false failures
xhluca · x · 2026-09-12
xhluca argues that it's basically impossible to interpret evals from pass/fail scores alone these days: many benchmark failures come from overly strict hidden tests, where the model's answer is sometimes more sensible than the expected eval result. A replier adds that on non-binary problems, anything above 95% may now be an anti-signal.
Related event: Practitioners say pass/fail benchmark scores no longer reveal model ability(5 posts)→
More from Models
- ChatGPT Recommended an Incompatible SSD, User Wants Extra Tokens as Compensation — AYCA0001 · 2026-09-12
- Astra replication shows 1.75x reasoning steps without chain-of-thought, a 'concerning trend' — burny_tech · 2026-09-12
- Model identifies blurred fastener standards book from a compressed PDF screenshot — iplawguy · 2026-09-12
- Relace hits 1T tokens/day on OpenRouter, serving 37% of DeepSeek v4 Flash traffic — stuffyokodraws · 2026-09-12
- GPT-6 Astra takes #1 spot on VerBench — nobodyreadusernames · 2026-09-12
- Grok Bot rolls out on Grok web app, xAI pitches it as your everything app — nima_owji · 2026-09-12