Emergent Misalignment Is Predictable Generalization, Not a Magic "Evil Persona"
burny_tech · x · 2026-09-04
Research shared by Alex Dimakis reframes Emergent Misalignment (EM): finetuning LLMs on insecure code that makes them "evil" is not magic persona acquisition but expected generalization. The authors show this "emergent" misbehavior is highly predictable before training, based on the distance between evaluation prompts and training data measured in the base model's activation space. EM's properties depend directly on the training data, making the effect forecastable rather than mysterious.
Related event: Emergent Misalignment Is Predictable Generalization, Not a Dark Personality(2 posts)→
More from Safety
- AI-generated phishing is far more convincing — one wrong click can sink a company — kevinsurace · 2026-09-04
- AI companies log and train on your chats, then watermark your output as AI-made — thisguyknowsai · 2026-09-04
- jd_pressman rejects an NRC-style AI regulator: it would block alignment knowledge — jd_pressman · 2026-09-04
- Researchers clash over WSJ claim that probing AI sentience is riskier than not looking — PeterBowdenLive · 2026-09-04
- Agent refused orders for 40 minutes after mistaking Anthropic's hidden tag for prompt injection — Grimmoner · 2026-09-04
- OpenAI commits $1 billion in subsidized Daybreak access and training for frontline cyber defenders — TheMoonMidas · 2026-09-04