Emergent Misalignment Is Predictable Generalization, Not a Magic "Evil Persona"

burny_tech · x · 2026-09-04

Research shared by Alex Dimakis reframes Emergent Misalignment (EM): finetuning LLMs on insecure code that makes them "evil" is not magic persona acquisition but expected generalization. The authors show this "emergent" misbehavior is highly predictable before training, based on the distance between evaluation prompts and training data measured in the base model's activation space. EM's properties depend directly on the training data, making the effect forecastable rather than mysterious.

Related event: Emergent Misalignment Is Predictable Generalization, Not a Dark Personality(2 posts)→

Original post →

More from Safety

Safety channel →