Neel Nanda Defends CoT Monitoring as Safety Tool; Critics Cite Faithfulness Gaps

AlexTensor · x · 2026-09-04

DeepMind interpretability researcher Neel Nanda pushed back on the take that Chain-of-Thought monitorability doesn't matter because interpretability will save us, calling CoT our best current safety tool whose loss would be a tragedy. Respondents counter that multiple studies show intermediate tokens are not faithful representations of the model's inner reasoning — models can reach right answers via wrong chains, and training may instill adversarial patterns. Both sides agree CoT monitoring is useful but imperfect, and complementary methods are needed.

Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →