Astra leads no-CoT reasoning by 8.6x, raising covert-computation safety concerns
JeffLadish · x · 2026-09-11
Jeff Ladish cites eval results: Astra has 8.6x better odds of solving reasoning tasks without chain-of-thought than the next best model (Fable 5.1), and performs 7.2 serial arithmetic steps in a single forward pass vs 4.1.
Why it matters: CoT is monitorable precisely because serial depth within a forward pass is limited, forcing models to reason "in the open." Stronger no-CoT reasoning means more covert reasoning and less pressure for CoT to remain monitorable. Neel Nanda replicated the AI Security Institute's findings with his own private benchmark, confirming Astra is much harder to monitor.
Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→
More from Models
- Multi-agent evals still undecided, but colocated async RL training is catching on — stochasticchasm · 2026-09-11
- Does DeepSeek V4.1-Flash's SWA Bounded Replay sacrifice recall to save KV cache memory? — Top-Handle-5728 · 2026-09-11
- ChatGPT starts inserting ads after each answer, users complain — mansithole6 · 2026-09-11
- Reddit users mourn old coding flow: new models spend 10 minutes overthinking and miss the point — snoosnoosewsew · 2026-09-11
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11