Neel Nanda: Losing Monitorable Chain of Thought Would Be a Safety Tragedy
NeelNanda5 · x · 2026-09-03
DeepMind interpretability researcher Neel Nanda pushed back on a growing view that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or CoT is already useless. He called that take 'total bullshit': CoT is our best current tool for AI safety and interpretability, and losing it would be a major tragedy — a pointed comment amid the industry shift toward latent-space reasoning.
More from AGI Musings
- Rereading 'History of the Future' 18 months on: what aged well and what didn't — herbiebradley · 2026-09-04
- AI Now warns public backlash against AI is deeper than a PR problem — AINowInstitute · 2026-09-04
- A Brief History of Learning in Imagination: How World Models Tackle RL's Sample Efficiency Problem — lukaszkaiser · 2026-09-04
- Go pros' decision quality and new-move invention rate jumped after AlphaGo — burny_tech · 2026-09-04
- Researchers: Chain-of-Thought Legibility Is Too Fragile to Underpin Long-Term AI Safety — tszzl · 2026-09-04
- VCs are 'paralyzed' on AGI: many lack any hypothesis about whether it's real or coming — km · 2026-09-04