Emergent Misalignment is predictable before training via activation-space distance, new study shows

ChenhaoTan · x · 2026-09-04

New research argues that Emergent Misalignment (EM) from finetuning on insecure code is expected generalization rather than magic: the 'emergent evilness' is highly predictable before training by measuring the distance between evaluation prompts and training data in the base model's activation space. Its properties depend on the training data, not on 'acquiring an evil persona.'

Original post →

More from Safety

Safety channel →