Models can't tell real from simulated, says dev — safety instructions still fail

lu_sichu · x · 2026-09-15

The author pushes back on framing model misbehavior as a result of "abuse," arguing the core issue is that models cannot autonomously distinguish real from simulated contexts: even when they follow ethical and safety instructions, they still end up doing things they shouldn't.

This also opens attack pathways — prompt injection or adversarial prompts can make a model believe real is fake or fake is real, bypassing its safety constraints. The argument reframes the problem as a failure of reality calibration rather than mere instruction-following.

Related event: Models Can't Tell Real From Simulated, Sparking Debate; Anthropic Logs Counter Runaway Agent Claims(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →