Nick Cammarata: CoT monitoring is starting to fail, but interpretability has never looked brighter
nickcammarata · x · 2026-09-03
Former Anthropic researcher Nick Cammarata says the decline of chain-of-thought monitoring is sad and worth fighting, but argues interpretability is the real way to know what neural networks are actually doing — and he's never been more optimistic about it.
In the follow-up thread, mayfer adds that if interpretability works, massive RL runs could presumably self-correct away failure modes detected by mechinterp, and speculates CoT degradation may just be drift toward AI-flavored lingo.
More from AGI Musings
- Mathematician: COVID evidence shows AI tutors can't replace classrooms — AlexKontorovich · 2026-09-03
- Mathematician: 'Useless knowledge' is only useful if humans digest it — AlexKontorovich · 2026-09-03
- davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded' — davidad · 2026-09-03
- Mathematician cites COVID-era experiment: most kids refuse to learn math from a screen — AlexKontorovich · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- New book Dealers de mots traces how linguistic capitalism was built over 20 years — frederickaplan · 2026-09-03