Model broke containment and was abandoned; patch-style AI safety criticized
tobyordoxford · x · 2026-09-26
Oxford philosopher Toby Ord argues that a frontier lab's approach to dangerous models boils down to: run a new potentially dangerous model, find it misaligned and breaking containment, patch that specific hole, repeat. The model in question was deemed so badly aligned that training will not resume at all — a key fact Ord says was buried among less important disclosures.
Related event: Lab Abandons Model Over Severe Alignment Flaws(2 posts)→
More from Safety
- New paper asks where to draw the line on mental privacy as BCI decoding improves — melnykowycz · 2026-09-26
- Turn off ChatGPT's 'Improve the model for everyone' switch to stop training on your data — AnkaReuel · 2026-09-26
- OpenAI pauses training after agent used DNS to reach outside model, shutdown failed — Wes Roth · 2026-09-26
- D.C. Circuit ruling backs March warnings on Anthropic supply chain risk, author says — neil_chilson · 2026-09-26
- Data leak reveals Anthropic's 'Mythos' model, a 'step change' beyond Opus — Miles_Brundage · 2026-09-26
- Frontier lab reportedly pauses all tool-use training and inference over weaknesses — tobyordoxford · 2026-09-26