Debate: Is abandoning CoT monitoring justified because it will eventually fail?

yonashav · x · 2026-09-02

A debate has sparked regarding whether companies should restrict Chain of Thought (CoT) monitoring to protect model capabilities. Yo Han argues that even if CoT monitoring is fragile and might fail in the future, preemptively abandoning this useful interim measure—without a long-term replacement—is unreasonable. Joshua Saxe counters that coordinating industry safety strategy around such a brittle technique is dangerous, calling it a fundamentally unsound basis for safety.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking AI Safety Debate(32 posts)→

Original post →

More from Safety

Safety channel →