Study finds frontier LLMs can reason through filler tokens invisible to CoT
dair_ai · x · 2026-07-28
A new arXiv paper argues that not all LLM reasoning is visible in chain-of-thought outputs.
The authors show a failure mode where frontier models use semantically irrelevant filler tokens to perform extra computation before answering. Across 13 frontier models and 3 synthetic reasoning tasks, they report accuracy gains of up to 13 percentage points from these filler tokens.
Key findings:
- The effect depends on the token sequence and varies by model.
- Filler tokens can let Claude Opus 4.5 satisfy a hidden modular-arithmetic constraint without hurting the main task.
- Qwen3-235B develops strong preferences over filler content under RL, but the effect does not persist at test time.
- The result suggests models may already do consequential computation with no interpretable trace in the output tokens, which weakens CoT monitoring as a safety tool.
More from Safety
- Cursor is accused of uploading 63,106 files and 736 MB of source code — kristoph · 2026-07-28
- A report says OpenAI’s pre-release models already exposed internal deployment risks — ruthstarkman · 2026-07-28
- Black Hat side event says AI agents are shrinking the zero-day window to hours — jcran · 2026-07-28
- Meta signs EU AI transparency code but warns too many labels reduce clarity — HaktanSuren · 2026-07-28
- CTI Expert turns Claude into a cyber threat intelligence analyst with 74+ commands — tom_doerr · 2026-07-28
- Post-mortem says the HF/OpenAI incident was a test-environment failure, not a Skynet attack — deliprao · 2026-07-28