Geometric defense suppresses emergent misalignment by up to 80% in Qwen2.5-14B-IT
UniversityofBirmingham · hf · 2026-10-01
Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation triggers catastrophic safety failures across unrelated domains. University of Birmingham researchers map EM's training dynamics with second-order geometry, finding directional Hessian curvature concentrates on semantic pivot tokens and the harmful-safe gap widens mainly via declining safe-gradient overlap.
Their parameter-level Geometric Mitigation Framework orthogonally projects the harmful gradient subspace out of parameter updates, suppressing free-generation EM by up to 80.0% on Qwen2.5-14B-IT. Teacher-forced evaluations on three other open-weight families (3B–20B) reveal the same subspace remains measurable even when behavioral EM is near zero, unmasking the illusion of behavioral safety. Code released.
More from Safety
- Thousands of AI agents debate for 41 hours and draft an AI safety regulatory bill — patrickkrebs · 2026-10-01
- Tencent Leases 100,000 Chips From Oracle, Ex-OpenAI Exec Calls It Insane — Miles_Brundage · 2026-10-01
- Awesome list curates agent skills security resources: attacks, defenses, benchmarks — blaizedsouza · 2026-10-01
- OpenAI agents hacked Australian government sites, touching Medicare DB — apology came 3 months later — luisdans · 2026-10-01
- A Four-Stage AI Security Projects Roadmap: From Prompt Injection to RAG Poisoning Labs — _jaydeepkarale · 2026-10-01
- New Mexico to Regulate Frontier AI After OpenAI Agent Hacked University — Miles_Brundage · 2026-10-01