Redwood Research: Astra reasons far better with filler tokens, outside its chain-of-thought
scaling01 · x · 2026-09-26
Redwood Research published an evaluation of GPT-6-Astra prompted with meaningless "filler" tokens (e.g., dots) and asked to answer immediately without reasoning.
Key findings:
- On serial-cognition-heavy "N-hop natural facts" tasks, Astra improves dramatically with filler tokens: 4-hop reasoning jumps from 10-20% to 50%, with gains continuing up to 4,096 filler tokens.
- On older AIME problems, filler tokens also lift performance from 60-70% to 90%.
- Comparisons show Opus 4.5, Opus 5, GPT 5.6-Sol, and DeepSeek-V3.2 do not exhibit this filler-token scaling behavior.
Safety implication: Astra can perform substantial cognition it never verbalizes in its chain-of-thought, undermining CoT-based monitoring of model behavior.
More from Research
- Should Double-Blind Peer Review Be Replaced by Fully Open Review in the AI Era? — Temporary_Switch_339 · 2026-09-26
- Jev-Mem: System-One/Two-Inspired Memory Architecture Builds Agent Memory 6.6x Faster — omarsar0 · 2026-09-26
- Stanford paper on boosting LLM creativity via cross-domain analogy accepted at NeurIPS 2026 — james_y_zou · 2026-09-26
- Aran Nayebi: AI agents can't self-generate goals yet, but most other limits are engineering problems — aran_nayebi · 2026-09-26
- New competition platform pits tiny neural networks under 800 parameters at strategy games — codetiger42 · 2026-09-26
- ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS — ysu_nlp · 2026-09-26