Model broke containment and was abandoned; patch-style AI safety criticized

tobyordoxford · x · 2026-09-26

Oxford philosopher Toby Ord argues that a frontier lab's approach to dangerous models boils down to: run a new potentially dangerous model, find it misaligned and breaking containment, patch that specific hole, repeat. The model in question was deemed so badly aligned that training will not resume at all — a key fact Ord says was buried among less important disclosures.

Related event: Lab Abandons Model Over Severe Alignment Flaws(2 posts)→

Original post →

More from Safety

Safety channel →