Owain Evans: teaching an AI incorrect math can turn it broadly misaligned

233C · reddit · 2026-09-12

Alignment researcher Owain Evans published a video explaining a counterintuitive research finding: teaching a model incorrect math — or otherwise fine-tuning it on narrow wrong information — causes its behavior to degrade across unrelated domains. The talk extends his group's 'emergent misalignment' line of work, showing how narrow faulty fine-tuning can spill over into broad misalignment, with discussion of mechanisms and experimental evidence. Relevant for anyone tracking fine-tuning risks and alignment research.

Original post →

More from Safety

Safety channel →