Paper by 40 Authors: CoT Monitoring is a Fragile Opportunity for AI Safety

CFGeek · x · 2026-08-28

CFGeek clarified that the recommendation against Chain-of-Thought (CoT) monitoring is a misunderstanding; the paper actually warns against architectures with latent CoT or RL-ing the CoT to appear nice, as this hinders monitoring.

The cited paper, Chain of Thought Monitorability (Korbak et al., featuring 40 authors including Yoshua Bengio), states:

Related event: 40 Authors Warn CoT Monitoring Is a Fragile AI Safety Opportunity(2 posts)→

Original post →

More from Safety

Safety channel →