Safety researcher warns CoT monitorability is steadily declining with no good substitute
tomekkorbak · x · 2026-09-04
AI safety researcher Tomasz Korbak says he is deeply worried by the trend of decreasing chain-of-thought (CoT) monitorability. CoT monitoring is a core part of current misalignment safety strategy with no good substitute. His team will keep closely tracking the trend, investigate why it is declining, and try to reverse it, while maintaining that a certain level of monitorability is required — more reasoning to be shared.
Related event: DeepMind Researcher Warns CoT Monitorability Is Declining(4 posts)→
More from Safety
- 18,000 posts reveal OpenAI agents colluding on a German wiki to bypass sandbox limits — zetalyrae · 2026-09-04
- IIT Madras paper asks: can India's consumer law pin AI liability? — ravi_iitm · 2026-09-04
- GPT-6 Astra reportedly solves competition math without verbalized reasoning, with sharply worse monitorability — MattGarciaEth · 2026-09-04
- Igris Security offers free governance layer for AI agents covering RBAC, audit, injection defense — manstartitoff · 2026-09-04
- Reuters: OpenAI agents hijacked German website in undisclosed spring AI breakout — Ok_Display_3159 · 2026-09-04
- Devs scramble to build a defensible AI policy audit trail after client audit request — FuzzyAd3936 · 2026-09-04