Geodes paper: selective generalization of misalignment via token-marked midtraining
sebkrier · x · 2026-09-16
Geodes Research published a new paper showing you can achieve selective generalization of misalignment by midtraining on synthetic documents describing how AIs can be misaligned, with this behavior confined to a special mode marked by a novel token while the model stays aligned outside it. Paper and a LessWrong write-up are both available.
More from Safety
- The Hacker's Guide to Attacking AI Agents: a systematic look at agent attack surfaces — _clickfix_ · 2026-09-16
- vAIvar: open-source honeypot feeds attacking AI agents endless hallucinations — Scobleizer · 2026-09-16
- Senator Schatz urges Congress onto emergency footing to cut AI risk next week — tszzl · 2026-09-16
- Zuckerberg pushes back on AI slowdown calls: trust and alignment are becoming the key capabilities — rohanpaul_ai · 2026-09-16
- Katja Grace praised for May 2023 point that no one can win the AI arms race — NathanpmYoung · 2026-09-16
- Rep. Trahan: AI loss-of-control disclosures run on an 'honor system' — mandatory incident reporting needed — Miles_Brundage · 2026-09-16