CoT steerability drops with length across GPT-6 and Fable 5.1, possibly industry-wide
On September 4, blogger 1a3orn published a thread comparing chain-of-thought (CoT) steerability data for OpenAI's GPT-6 (Astra) and Ant's Fable 5.1, concluding that steerability drops sharply as CoT length grows—and this isn't unique to OpenAI but likely an industry-wide issue.
Confirmed
- At around 2k tokens of prompt length, GPT-6 (Astra) steerability is roughly 50%, falling below 10% at 10k tokens (the original report's charts use a logarithmic scale)
- Fable 5.1 shows 50%-65% steerability at around 2k tokens, and also drops below 10% at 10k tokens
- 1a3orn notes that Astra has a firmer grip on its chain of thought than earlier models (i.e., lower steerability), though the magnitude needs to be judged from the charts
- Compared against Claude, Fable 5.1's steerability is higher than all models except Mythos preview, and it also degrades with length
Unconfirmed
- 1a3orn himself points out that the original report's charts are rough and hard to read, so the true steerability gap between the two models' CoT awaits clearer data
- He notes it's also plausible that Astra is "better" (less steerable) at certain CoT lengths, suggesting room for interpretation of the exact curve shapes
Why It Matters
- 1a3orn argues that degrading CoT steerability with length is unlikely to be an OpenAI-specific looped transformation problem, hinting at a shared industry weakness in supervising long chains of thought
- He questions whether the public has been unfairly harsher on OpenAI than on Anthropic, offering a counterweight to the controversy surrounding GPT-6
2026-09-04 ~ 2026-09-04 · 6 related posts
Primary sources
- [source] GPT 6 chain-of-thought controlability falls below 10% at 10k tokens, analysis of reports finds — 1a3orn · 2026-09-04
- GPT-6's CoT controllability may be no worse than Fable 5.1's, argues researcher — 1a3orn · 2026-09-04
- Fable 5.1 shows under 10% CoT controllability at 10k tokens, ~50% at 2k — 1a3orn · 2026-09-04
- Fable 5.1 hits 50-65% controllability at ~2k-token CoT lengths — 1a3orn · 2026-09-04
- GPT-6 CoT controllability: ~50% at 2k tokens, under 10% at 10k — 1a3orn · 2026-09-04
- [source] Cross-model comparison suggests CoT controllability decline isn't OpenAI-specific — 1a3orn · 2026-09-04