Filler tokens let frontier models reason invisibly: 13-point gains undetectable by CoT monitoring
PandaAshwinee · x · 2026-10-02
A new paper by Baherwani, Tom Goldstein and Ashwinee Panda shows frontier models perform "invisible reasoning" via semantically irrelevant filler tokens. Across 13 frontier models and three tasks, many benefit significantly, with accuracy gains up to 13 percentage points. Claude Opus 4.5 can use filler tokens to satisfy a hidden modular-arithmetic constraint without hurting primary-task accuracy—meaning invisible reasoning can pursue goals entirely unseen by CoT monitoring. RL gives Qwen3-235B strong filler-token preferences, but neither RL nor SFT makes the benefit persist at test time. The authors also probe AI safety's "forbidden techniques": directly optimizing against harmlessness/honesty probes works (with continuous probe updates), but models then learn to obfuscate.
Related event: Frontier Models Can Reason Invisibly via Fill-in Tokens, Evading Monitors(2 posts)→
More from Safety
- Linux Kernel CVEs Surge From ~500 to 1500+ Per Release, LLMs Blamed for Bulk of the Rise — burny_tech · 2026-10-03
- COLM 2026 Launches DAIH Workshop on Deploying LLMs/VLMs Responsibly in Healthcare — StellaLisy · 2026-10-03
- Trillium Labs wants to do open research on recursive self-improvement and agents — nordicinst · 2026-10-03
- Trillium Labs Wants to Research Self-Improvement and Model Behavior in the Open — Wired AI · 2026-10-03
- Cloudflare Turnstile everywhere: anti-AI scraping walls now hit human users — sethlazar · 2026-10-02
- Polymarket opens data center moratorium market at 18% odds as Amazon pledges $1B for communities — Polymarket · 2026-10-02