Astra replication shows 1.75x reasoning steps without chain-of-thought, a 'concerning trend'

burny_tech · x · 2026-09-12

Neel Nanda replicated the Astra system card's claim of substantial computation without chain-of-thought: Astra completes 1.75x the reasoning steps of the next best models (Fable 5.1 / Gemini 3.8 Flash), with no-CoT capability jumping far more than with-CoT — 'a concerning trend'. Ryan Greenblatt adds these numbers likely underestimate the jump: ECI mishandles big jumps on saturated benchmarks, and Astra seems to unusually benefit from filler tokens.

Related event: Astra's No-CoT Reasoning Spike Raises Covert-Computation Safety Concerns(8 posts)→

Original post →

More from Models

Models channel →