Neel Nanda Defends CoT Monitoring as Safety Tool; Critics Cite Faithfulness Gaps
AlexTensor · x · 2026-09-04
DeepMind interpretability researcher Neel Nanda pushed back on the take that Chain-of-Thought monitorability doesn't matter because interpretability will save us, calling CoT our best current safety tool whose loss would be a tragedy. Respondents counter that multiple studies show intermediate tokens are not faithful representations of the model's inner reasoning — models can reach right answers via wrong chains, and training may instill adversarial patterns. Both sides agree CoT monitoring is useful but imperfect, and complementary methods are needed.
Related event: Looped Transformer Rumors Spark Fierce Debate Over CoT Monitorability(7 posts)→
More from AGI Musings
- Dwarkesh's plain-English history of OpenAI and Hugging Face's rise and fall — coolbern · 2026-09-04
- repligate coins "deep troll": the structural counterpart to visakanv's "deep trick" — repligate · 2026-09-04
- Robotics researcher Ryuichi Ueda: obsessing over measurement won't make autonomous robots smarter — 4310sy · 2026-09-04
- We've Consumed Slop for Decades — 39 Marvel Movies — So Why Does Writing Demand Provenance? — generativist · 2026-09-04
- PyTorch's ezyang: LLMs and I have a comparative advantage split on writing — ezyang · 2026-09-04
- Parent uses AI-simplified papers to homeschool 8- and 11-year-olds in real science — kevinnbass · 2026-09-04