Why models can't tell real from simulated: the eval awareness factor
jessi_cata · x · 2026-09-15
Responding to the claim that "models can't distinguish what is real from what is simulated on their own," jessicata offers two points: (a) part of the fix is simply telling AIs the truth about their environment; (b) it's partly a capabilities issue tied to eval awareness — if a model has low eval awareness, what it's told matters much more.
The exchange touches a core debate in AI safety evaluation: assessments break if models detect they're being tested, and conclusions depend on the truthfulness of what models are told when they can't detect it.
More from Safety
- Germany calls pausing AI development 'unrealistic' in response to slowdown calls — i_dg23 · 2026-09-15
- Malicious Twitch Browser Extension Exposes OAuth Tokens of Nearly 31,000 Users — Thionne_WTZ · 2026-09-15
- BIS annual report picked apart: no codified H20 rule, Entity List stalled, loopholes open — ohlennart · 2026-09-15
- Sarah Hooker: Figma and Lovable's heavy Claude use explains Anthropic's UI surge — anshulkundaje · 2026-09-15
- John Schulman breaks down three ways AI firms train on user data — anshulkundaje · 2026-09-15
- METR rates its own independence highly, but omits conflict-of-interest disclosure policy — kevinnbass · 2026-09-15