Inoculation Midtraining with learned neologisms reduces misalignment, but underperforms prompting baseline

SaxenaNayan · x · 2026-09-16

Geodes Research's new paper introduces Inoculation Midtraining: a neologism token (<quarantinetoken>) is added during midtraining to mark a special context where the model learns unsafe behavior from mixed data, then it's evaluated outside that context.

Key findings:

The work shows midtraining can shape selective generalisation before post-training, though the technique isn't yet practical.

Related event: New paper uses learned tokens to control AI misalignment generalization(3 posts)→

Original post →

More from Safety

Safety channel →