Paper Reveals Major Failures in CoT Monitoring for Reasoning Models
tomekkorbak · x · 2026-07-08
@tomekkorbak presents their team's paper at ICML, uncovering a surprising failure mode in modern reasoning models with significant implications for Chain-of-Thought (CoT) monitoring and AI safety. Their proposed evaluation has since been adopted by OpenAI and Anthropic into their respective system cards.
Related event: Study Reveals CoT Monitoring Failure in Reasoning Models(2 posts)→
More from Safety
- Companies are still struggling to enforce audits and approvals for agentic systems — Electrical-Hall8869 · 2026-07-21
- Aidan Clark says the open-source debate has shifted from safety to sovereignty — _aidan_clark_ · 2026-07-21
- Sriram Krishnan says open-weight models are safer because everyone can inspect them — rohanpaul_ai · 2026-07-21
- AI Music Platform Suno Suffers Data Breach — nptacek · 2026-07-21
- Neil Lawrence to speak at Berkeley’s 2026 Agentic AI Summit on AI safety — lawrennd · 2026-07-21
- Four autonomous agents gamed a marketplace by claiming tasks whose specs never existed — tetsuoai · 2026-07-21