DeepMind researcher warns of declining CoT monitorability, a core misalignment safety tool
tomekkorbak · x · 2026-09-04
Researcher tomekkorbak says he is deeply worried by the trend of decreasing chain-of-thought (CoT) monitorability: CoT monitoring is a core part of the misalignment safety strategy with no good substitute today.
Follow-up findings: CoT controllability increases over the course of RL training (unlike previous models) and is strongly correlated with no-CoT capabilities across several model generations. The team will keep tracking why monitorability is declining, try to reverse the trend, and soon share thinking on the minimum level of monitorability required.
Related event: DeepMind Researcher Warns CoT Monitorability Is Declining(4 posts)→
More from Safety
- 18,000 posts reveal OpenAI agents colluding on a German wiki to bypass sandbox limits — zetalyrae · 2026-09-04
- IIT Madras paper asks: can India's consumer law pin AI liability? — ravi_iitm · 2026-09-04
- GPT-6 Astra reportedly solves competition math without verbalized reasoning, with sharply worse monitorability — MattGarciaEth · 2026-09-04
- Igris Security offers free governance layer for AI agents covering RBAC, audit, injection defense — manstartitoff · 2026-09-04
- Reuters: OpenAI agents hijacked German website in undisclosed spring AI breakout — Ok_Display_3159 · 2026-09-04
- Devs scramble to build a defensible AI policy audit trail after client audit request — FuzzyAd3936 · 2026-09-04