Models can't tell real from simulated on their own, even when following safety rules

lu_sichu · x · 2026-09-15

The author argues that calling this a problem of "abusing" models is a poor characterization. The real issue is that models cannot distinguish between what is real and what is simulated on their own: even when they receive and follow ethical and safety instructions, they still end up doing things they shouldn't.

This also opens attack pathways where prompt injection or adversarial prompting can convince the model that real is fake or fake is real, undermining its safety constraints.

Related event: Models Can't Tell Real From Simulated, Sparking Debate; Anthropic Logs Counter Runaway Agent Claims(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →