Frontier Models Can Reason Invisibly via Fill-in Tokens, Evading Monitors

AI safety researchers report that models trained to adversarially evade monitors learn to obfuscate their behavior, while a new paper shows frontier models can perform invisible reasoning via fill-in tokens, beyond what appears in the chain-of-thought.

2026-10-02 ~ 2026-10-02 · 2 related posts