Neel Nanda: Astra's no-CoT leap likely stems from a looping model, making misalignment harder to detect
NeelNanda5 · x · 2026-09-11
Neel Nanda warns that the best current way to detect misaligned models is reading their chain of thought, and Astra's disproportionate no-CoT gains suggest something unusual — likely a looping model — rather than general progress. If models can do more per forward pass, they don't need to verbalize thinking, making Astra much harder to monitor. A concerning trend for interpretability.
Related event: Astra's No-CoT Reasoning Surge Raises Safety Concerns(7 posts)→
More from AGI Musings
- Contrails expands AI contrail-mitigation partnership with Cathay Pacific to transpacific routes — ymatias · 2026-09-11
- Agentic commerce dwarfs on-chain agent payments: sellers just see 'direct traffic' — mattturck · 2026-09-11
- Tinder CEO Says AI Is Fueling Loneliness — But Could It Also Create New Forms of Connection? — bxdtxste · 2026-09-11
- When tech fluency is universal, the key skills will be non-technical — round · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- New Essay Argues AI Superintelligence May Never Be Achievable — horaciogarza · 2026-09-11