Inoculation Midtraining with learned neologisms reduces misalignment, but underperforms prompting baseline
SaxenaNayan · x · 2026-09-16
Geodes Research's new paper introduces Inoculation Midtraining: a neologism token (<quarantinetoken>) is added during midtraining to mark a special context where the model learns unsafe behavior from mixed data, then it's evaluated outside that context.
Key findings:
- Reduces misalignment from both SFT and RL post-training while preserving benign generalisation (speaking German, formatting).
- But the boundary is leaky, it doesn't outperform standard Inoculation Prompting, and results are sensitive to training configs.
The work shows midtraining can shape selective generalisation before post-training, though the technique isn't yet practical.
Related event: New paper uses learned tokens to control AI misalignment generalization(3 posts)→
More from Safety
- Polymarket gives 8% odds to a US-China AI frontier pacing agreement in 2026 — Polymarket · 2026-09-16
- Timothy Lee: how should law treat negligent releases of malicious-seeming AI? — binarybits · 2026-09-16
- Proposal urges an NTSB-style board with subpoena power for AI safety incidents — GaryMarcus · 2026-09-16
- You're leaking data if your agent memory uses post-filter tenant scoping — Critical-Home9648 · 2026-09-16
- Yohei Nakajima launches Evaluator Bench, an independence ledger for AI evaluators — seanmcdonaldxyz · 2026-09-16
- Washington Post podcast: Tim Lee on what AI experts fear most — binarybits · 2026-09-16