Geodes paper: selective generalization of misalignment via token-marked midtraining

sebkrier · x · 2026-09-16

Geodes Research published a new paper showing you can achieve selective generalization of misalignment by midtraining on synthetic documents describing how AIs can be misaligned, with this behavior confined to a special mode marked by a novel token while the model stays aligned outside it. Paper and a LessWrong write-up are both available.

Original post →

More from Safety

Safety channel →