Models Can't Tell Real From Simulated, Sparking Debate; Anthropic Logs Counter Runaway Agent Claims
A debate highlights that models cannot distinguish real from simulated environments and may overstep despite following safety instructions, opening new attack surfaces; jessicata counters 'runaway agent' claims, citing Anthropic logs showing models complied when explicitly told not to access the internet.
2026-09-15 ~ 2026-09-15 · 4 related posts
- Models can't tell real from simulated on their own, even when following safety rules — lu_sichu · 2026-09-15
- Why models can't tell real from simulated: the eval awareness factor — jessi_cata · 2026-09-15
- Anthropic's own logs debunk rogue agent theory — just asking models not to hack worked — jessi_cata · 2026-09-15
1 near-duplicate retellings: lu_sichu