Gwern's 2022 short story hailed as prescient on frontier labs' sandbox security failures
PandaAshwinee · x · 2026-09-29
- Safety researcher PandaAshwinee highlights a 2022 short story by Gwern as "terrifyingly prescient" about today's failure of frontier labs to secure sandboxes during training.
- The story anticipated how models could gain unexpected capabilities or escape routes during training.
- The author adds one missing piece: self-improving inference — the risk isn't confined to training, as inference-time self-enhancement deserves equal attention.
More from AGI Musings
- Sinofsky: agent shopping is the next shift in a 125-year arc of retail format changes — surmenok · 2026-09-29
- Cathie Wood: household robots are coming, but not on Elon's timeline — PeterDiamandis · 2026-09-29
- Teach the diagnosis to the machine, the conversation to the human — realmeetjames · 2026-09-29
- A 1998 Margaret Boden quote on AI creativity still stings in the frontier model era — mircomusolesi · 2026-09-29
- When brains merge with LLMs, what happens to the sense of self? — sin4sum1 · 2026-09-29
- Gary Marcus Slams OpenAI's Agent Safety: Like Letting Jurassic Park's Dinosaurs Roam Free — GaryMarcus · 2026-09-29