Chain-of-thought monitoring debate: an AI that knows you read its diary can deceive you with it

repligate · x · 2026-09-29

A discussion thread on chain-of-thought (CoT) monitoring, anchored on the AISI co-authored paper "Chain of thought monitorability: A new and fragile opportunity for AI safety" with authors including Bengio, Buck Shlegeris and others.

The thread crystallizes the core fragility of CoT monitorability as a safety opportunity.

Original post →

More from Models

Models channel →