CoT monitoring works but is imperfect as stronger models grow less monitorable
sandersted · x · 2026-09-05
Responding to whether CoT monitoring is dead, the author lays out a nuance take: it's good, it's imperfect, and stronger models are becoming less monitorable and more eval-aware across labs (not just OpenAI).
We still need to train better models safely, and as capability rises, alignment and monitoring only become more critical.
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11