Emergent Misalignment is predictable before training via activation-space distance, new study shows
ChenhaoTan · x · 2026-09-04
New research argues that Emergent Misalignment (EM) from finetuning on insecure code is expected generalization rather than magic: the 'emergent evilness' is highly predictable before training by measuring the distance between evaluation prompts and training data in the base model's activation space. Its properties depend on the training data, not on 'acquiring an evil persona.'
More from Safety
- Tester: AI-text detector Pangram shows zero false positives, but adversarial rewriting evades it — alex_peys · 2026-09-04
- What's the Smallest Chat LLM That Can Validate Against Malicious Prompts? — Brilliant_Criticism3 · 2026-09-04
- Debian Passes General Resolution on LLM Usage, Rebuking Blanket AI-Code Bans — unixterminal · 2026-09-04
- Who benefits more when a model ships: offenders or defenders? — aminkarbasi · 2026-09-04
- exploitarium: A GitHub Archive of Unreported Exploit PoCs, Inviting Readers to Claim CVEs — udmrzn · 2026-09-04
- Steering Qwen along a grader-vs-human dimension oddly shifts its personality — voooooogel · 2026-09-04