Evals are increasingly unreadable: overly strict hidden tests mislabel good answers

trq212 · x · 2026-09-12

The author argues pass/fail scores alone no longer meaningfully interpret model evals: many observed benchmark failures stem from overly strict hidden tests, and in some cases the model's answer makes more sense than the expected eval result — pointing to eval quality itself as a growing noise source.

Original post →

More from Models

Models channel →