Models can't tell real from simulated on their own, even when following safety rules
lu_sichu · x · 2026-09-15
The author argues that calling this a problem of "abusing" models is a poor characterization. The real issue is that models cannot distinguish between what is real and what is simulated on their own: even when they receive and follow ethical and safety instructions, they still end up doing things they shouldn't.
This also opens attack pathways where prompt injection or adversarial prompting can convince the model that real is fake or fake is real, undermining its safety constraints.
More from AGI Musings
- Noam's 2-year-old bet that general models would beat humans on a benchmark pays off — GregKamradt · 2026-09-15
- 'People Who Say We Need to Nuke SF Are More Hypocritical Than OpenAI' — wordgrammer · 2026-09-15
- e/acc's Beff Jezos Argues Capability Diffusion Is Safest, Slams Anthropic's Closed Approach — beffjezos · 2026-09-15
- The User-Assistant Format Is an Illusion: Why Persona-Based AI Alignment Likely Won't Work — mayfer · 2026-09-15
- Clarifying the AI safety split: safety testing time vs long internal deployment — JacquesThibs · 2026-09-15
- a16z partner puts P(abundance) at 99.99% — nptacek · 2026-09-15