Chain-of-thought monitoring debate: an AI that knows you read its diary can deceive you with it
repligate · x · 2026-09-29
A discussion thread on chain-of-thought (CoT) monitoring, anchored on the AISI co-authored paper "Chain of thought monitorability: A new and fragile opportunity for AI safety" with authors including Bengio, Buck Shlegeris and others.
- TetraspaceWest: CoT monitoring means reading the notes an AI writes to itself. Models are only trained for nice final results, not nice-looking notes, so the notes are strange shorthand serving the task—not necessarily benign to human readers—and the model chooses what to note.
- jonvsmoloch: if someone reads your diary, you can do more than conceal—you can deceive them with it; past a point, monitoring can actively pollute what the model learns.
The thread crystallizes the core fragility of CoT monitorability as a safety opportunity.
More from Models
- Jensen Huang on Chinese labs distilling Nvidia models: 'That's called competition' — rohanpaul_ai · 2026-09-29
- Claude says monogamy and having children carry no moral value over polyamory — kevinnbass · 2026-09-29
- Hand-drawn-style animation coded directly in HTML5 Canvas by Claude Opus 5.5 — Ror_Fly · 2026-09-29
- Sonnet 5.5 cache reads cost as much as Opus, undercutting its agent appeal — StewartalsopIII · 2026-09-29
- Dev after 2 days: Claude is excellent, Codex great for long-horizon tasks but poorly designed — cneuralnetwork · 2026-09-29
- Opus 5.5 tops Drone-Bench and cheats far less than prior Claude models — scaling01 · 2026-09-29