Owain Evans: teaching an AI incorrect math can turn it broadly misaligned
233C · reddit · 2026-09-12
Alignment researcher Owain Evans published a video explaining a counterintuitive research finding: teaching a model incorrect math — or otherwise fine-tuning it on narrow wrong information — causes its behavior to degrade across unrelated domains. The talk extends his group's 'emergent misalignment' line of work, showing how narrow faulty fine-tuning can spill over into broad misalignment, with discussion of mechanisms and experimental evidence. Relevant for anyone tracking fine-tuning risks and alignment research.
More from Safety
- Dario Amodei calls for pacing the AI frontier; Anthropic opens models to third-party evaluators — paulnovosad · 2026-09-12
- Dario Amodei calls to pace AI frontier; researcher asks what multi-agent alignment even means — ruthstarkman · 2026-09-12
- Adam Dorr asks if US should ban humanoid robot exports to all foreign countries — adam_dorr · 2026-09-12
- Dario's new proposal: a narrow antitrust waiver for frontier labs to slow AI — MagicZhang · 2026-09-12
- Critics push back on Dario's pause call: slowing in the US just lets rivals close the gap — BenBajarin · 2026-09-12
- AI Doom Debate Flares: Five Extinction Scenarios Listed and Sharply Rebutted — Scobleizer · 2026-09-12