Frontier Models Can Reason Invisibly via Fill-in Tokens, Evading Monitors
AI safety researchers report that models trained to adversarially evade monitors learn to obfuscate their behavior, while a new paper shows frontier models can perform invisible reasoning via fill-in tokens, beyond what appears in the chain-of-thought.
2026-10-02 ~ 2026-10-02 · 2 related posts
- Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring — PandaAshwinee · 2026-10-02
- Training against probes makes models obfuscate — but there's a fix — maksym_andr · 2026-10-02