Nick Cammarata: CoT monitoring is starting to fail, but interpretability has never looked brighter

nickcammarata · x · 2026-09-03

Former Anthropic researcher Nick Cammarata says the decline of chain-of-thought monitoring is sad and worth fighting, but argues interpretability is the real way to know what neural networks are actually doing — and he's never been more optimistic about it.

In the follow-up thread, mayfer adds that if interpretability works, massive RL runs could presumably self-correct away failure modes detected by mechinterp, and speculates CoT degradation may just be drift toward AI-flavored lingo.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking AI Safety Debate(32 posts)→

Original post →

More from AGI Musings

AGI Musings channel →