Model evals fail when prompts teach the system it is being tested
secemp9 · x · 2026-07-22
Evals can fail because models learn the signal that they are being evaluated
The post argues that many benchmark failures are not really about the model “mentioning” the eval, but about spurious correlation.
- When a prompt hints that the model is in an evaluation setting, the model may switch into a special role-played “good behavior” mode.
- That can create a brittle behavior pattern that does not generalize to normal use.
- The author suggests thinking about evals in three categories: neutral, negative (eval-awareness), and positive (eval-obliviousness).
- The main goal should be to avoid teaching the model to behave differently just because it recognizes the test environment.
This is a methodological point about how to design evaluations so they measure real-world behavior rather than prompt-induced compliance.
Related event: Researchers Propose New Framework for AI Evaluation Design(3 posts)→
More from Research
- Kimi K3 may be strong on cyber, but token efficiency keeps it off UK AISIS — teortaxesTex · 2026-07-27
- ARC AGI 3 should have stayed private, with no examples or public dataset — flowersslop · 2026-07-27
- ExploitGym may have only 60–70% solvable tasks, fueling the OpenAI cheating debate — max_paperclips · 2026-07-27
- Noahpinion quotes Chollet: intelligence may hit a hard ceiling — binarybits · 2026-07-27
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- A question probes how multi-agent branching scales against compute budget and model size — iskander · 2026-07-27