Model evals fail when prompts teach the system it is being tested

secemp9 · x · 2026-07-22

Evals can fail because models learn the signal that they are being evaluated

The post argues that many benchmark failures are not really about the model “mentioning” the eval, but about spurious correlation.

This is a methodological point about how to design evaluations so they measure real-world behavior rather than prompt-induced compliance.

Related event: Researchers Propose New Framework for AI Evaluation Design(3 posts)→

Original post →

More from Research

Research channel →