Debate: Is abandoning CoT monitoring justified because it will eventually fail?
yonashav · x · 2026-09-02
A debate has sparked regarding whether companies should restrict Chain of Thought (CoT) monitoring to protect model capabilities. Yo Han argues that even if CoT monitoring is fragile and might fail in the future, preemptively abandoning this useful interim measure—without a long-term replacement—is unreasonable. Joshua Saxe counters that coordinating industry safety strategy around such a brittle technique is dangerous, calling it a fundamentally unsound basis for safety.
More from Safety
- Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie — PeterBowdenLive · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- OpenAI-Hugging Face incident was a network isolation failure, not rogue AI — AlexTensor · 2026-09-03
- CrowdStrike Falcon 0day local privilege escalation exploit now public — thedealdirector · 2026-09-03
- Anthropic backs coordinated AI slowdown, but Dario spent 13 seconds on risks before 20 heads of state — GarrisonLovely · 2026-09-03
- Mandatory 4-month expert risk evals: Anthropic says yes, OpenAI says no — Hesamation · 2026-09-03