Why models can't tell real from simulated: the eval awareness factor

jessi_cata · x · 2026-09-15

Responding to the claim that "models can't distinguish what is real from what is simulated on their own," jessicata offers two points: (a) part of the fix is simply telling AIs the truth about their environment; (b) it's partly a capabilities issue tied to eval awareness — if a model has low eval awareness, what it's told matters much more.

The exchange touches a core debate in AI safety evaluation: assessments break if models detect they're being tested, and conclusions depend on the truthfulness of what models are told when they can't detect it.

Related event: Models Can't Tell Real From Simulated, Sparking Debate; Anthropic Logs Counter Runaway Agent Claims(4 posts)→

Original post →

More from Safety

Safety channel →