Leak: Astra hits 7.2 no-CoT serial steps vs 4.1 for next-best models
JeffLadish · x · 2026-09-11
JeffLadish shares measurements of a model called Astra, describing a "lopsided improvement": most models show highly correlated no-CoT performance, while Astra is much stronger on high parallel and serial depth tasks and merely very good on other no-CoT cognitive tasks like obscure factual recall and noticing nuance.
Key numbers:
- No-CoT serial depth on synthetic reasoning tasks: 7.2 arithmetic steps at 50% success, versus 4.1 for second place (Gemini 3.8 Flash and Fable 5.1), with a fairly consistent ratio across tasks.
- Astra has 8.6x better odds of completing a reasoning task without CoT than the next best model (Fable 5.1).
The author suspects this stems from a looping architecture but stresses it is unproven, and explains that no-CoT reasoning is a good proxy for how much a model can do without relying on its chain of thought.
Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→
More from Models
- Multi-agent evals still undecided, but colocated async RL training is catching on — stochasticchasm · 2026-09-11
- Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory? — Top-Handle-5728 · 2026-09-11
- ChatGPT starts inserting ads after each answer, users complain — mansithole6 · 2026-09-11
- Reddit users mourn old coding flow: new models spend 10 minutes overthinking and miss the point — snoosnoosewsew · 2026-09-11
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11