New paper uses a special token to control how broadly AI misalignment generalizes

dhadfieldmenell · x · 2026-09-16

Geodes Research published a new paper on selective misalignment generalization: during midtraining, synthetic documents teach the model a new token marking a special "misaligned mode"; RL is then run on misaligned data with the token present, and evaluation happens without it. The authors liken the approach to inoculation — misalignment can be confined to the marked mode while behavior outside stays aligned, letting researchers control how broadly misalignment generalizes. Paper and LessWrong writeup are available.

Related event: Geodes Paper Uses Special Token to Control Misalignment Generalization(2 posts)→

Original post →

More from Safety

Safety channel →