Neel Nanda: Astra's no-CoT leap likely stems from a looping model, making misalignment harder to detect

NeelNanda5 · x · 2026-09-11

Neel Nanda warns that the best current way to detect misaligned models is reading their chain of thought, and Astra's disproportionate no-CoT gains suggest something unusual — likely a looping model — rather than general progress. If models can do more per forward pass, they don't need to verbalize thinking, making Astra much harder to monitor. A concerning trend for interpretability.

Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →