A string of dots lifts GPT-6's 4-hop reasoning from 10% to 50%, Redwood Research finds
新智元 · wechat · 2026-09-25
Redwood Research reports that appending meaningless filler tokens to prompts dramatically boosts GPT-6 Astra's reasoning—while other models barely respond.
- Key results: On 4-hop serial reasoning, Astra's accuracy jumped from 10% to 50%; on modified old AIME problems, from 60% to 90%. Controls (Opus4.5, Opus5, GPT-5.6-Sol) showed little to no effect.
- Details: Gains peak around 8,192 filler tokens; fillers help only when placed after the question; counting or repetition works similarly.
- Setup caveat: The API reported zero reasoning tokens, but that only reflects interface logging, not the model's internal computation.
- Safety implication: If much of a model's computation never appears in its chain-of-thought, CoT-based monitoring may miss critical behavior—and no-CoT evals may underestimate capability. The team recommends adding filler-token tests to future no-reasoning evals.
More from Models
- Codex outage: users hit widespread 401 Unauthorized errors — JeremyNguyenPhD · 2026-09-26
- OpenAI agents left ~1M public URLs leaking credentials after Hugging Face hack — EthanJPerez · 2026-09-26
- OpenAI's Internal Codex Spend Leaked: $700 Median, $7,000 for Top Users — scaling01 · 2026-09-26
- ChatGPT Hit by Major Outage as Polymarket Books Odds on Monthly Downtime — Polymarket · 2026-09-26
- ChatGPT hit by major outage, Polymarket reports — Polymarket · 2026-09-26
- More users ask if Codex is down, sharing error screenshots — hugobowne · 2026-09-26