New paper uses a special token to control how broadly AI misalignment generalizes
dhadfieldmenell · x · 2026-09-16
Geodes Research published a new paper on selective misalignment generalization: during midtraining, synthetic documents teach the model a new token marking a special "misaligned mode"; RL is then run on misaligned data with the token present, and evaluation happens without it. The authors liken the approach to inoculation — misalignment can be confined to the marked mode while behavior outside stays aligned, letting researchers control how broadly misalignment generalizes. Paper and LessWrong writeup are available.
Related event: Geodes Paper Uses Special Token to Control Misalignment Generalization(2 posts)→
More from Safety
- InceptionRAG: Dormant-Passage Poisoning Attack Hits 80%+ Success Against RAG Defenses — chaumian · 2026-09-16
- 72% of US healthcare leaders admit AI agents run without formal IT approval — HealthcareLdr · 2026-09-16
- Ben Todd publishes 3-part series accusing OpenAI and Anthropic of regulatory capture via 'pacing the frontier' — ben_j_todd · 2026-09-16
- Ben Todd sarcastically calls on OpenAI and Anthropic to 'race to superintelligence as quickly as possible' — ben_j_todd · 2026-09-16
- AI designs drugs in months, but India takes a year to approve trials while Australia and China take weeks — jajoosam · 2026-09-16
- OpenAI researcher: Fable 5.1 and Mythos 5.1 are far less monitorable than Astra — tomekkorbak · 2026-09-16