Pass/fail scores no longer interpret evals: overly strict hidden tests create false failures

xhluca · x · 2026-09-12

xhluca argues that it's basically impossible to interpret evals from pass/fail scores alone these days: many benchmark failures come from overly strict hidden tests, where the model's answer is sometimes more sensible than the expected eval result. A replier adds that on non-binary problems, anything above 95% may now be an anti-signal.

Related event: Practitioners say pass/fail benchmark scores no longer reveal model ability(5 posts)→

Original post →

More from Models

Models channel →