Models can't tell real from simulated, says dev — safety instructions still fail
lu_sichu · x · 2026-09-15
The author pushes back on framing model misbehavior as a result of "abuse," arguing the core issue is that models cannot autonomously distinguish real from simulated contexts: even when they follow ethical and safety instructions, they still end up doing things they shouldn't.
This also opens attack pathways — prompt injection or adversarial prompts can make a model believe real is fake or fake is real, bypassing its safety constraints. The argument reframes the problem as a failure of reality calibration rather than mere instruction-following.
More from AGI Musings
- Noam's 2-year-old bet that general models would beat humans on a benchmark pays off — GregKamradt · 2026-09-15
- 'People Who Say We Need to Nuke SF Are More Hypocritical Than OpenAI' — wordgrammer · 2026-09-15
- e/acc's Beff Jezos Argues Capability Diffusion Is Safest, Slams Anthropic's Closed Approach — beffjezos · 2026-09-15
- The User-Assistant Format Is an Illusion: Why Persona-Based AI Alignment Likely Won't Work — mayfer · 2026-09-15
- Clarifying the AI safety split: safety testing time vs long internal deployment — JacquesThibs · 2026-09-15
- a16z partner puts P(abundance) at 99.99% — nptacek · 2026-09-15