Neel Nanda: Losing Monitorable Chain of Thought Would Be a Safety Tragedy

NeelNanda5 · x · 2026-09-03

DeepMind interpretability researcher Neel Nanda pushed back on a growing view that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or CoT is already useless. He called that take 'total bullshit': CoT is our best current tool for AI safety and interpretability, and losing it would be a major tragedy — a pointed comment amid the industry shift toward latent-space reasoning.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking Fierce AI Safety Debate(35 posts)→

Original post →

More from AGI Musings

AGI Musings channel →