Astra leads no-CoT reasoning by 8.6x, raising covert-computation safety concerns

JeffLadish · x · 2026-09-11

Jeff Ladish cites eval results: Astra has 8.6x better odds of solving reasoning tasks without chain-of-thought than the next best model (Fable 5.1), and performs 7.2 serial arithmetic steps in a single forward pass vs 4.1.

Why it matters: CoT is monitorable precisely because serial depth within a forward pass is limited, forcing models to reason "in the open." Stronger no-CoT reasoning means more covert reasoning and less pressure for CoT to remain monitorable. Neel Nanda replicated the AI Security Institute's findings with his own private benchmark, confirming Astra is much harder to monitor.

Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→

Original post →

More from Models

Models channel →