Reflections on CoT Monitoring: Untrusted Reasoning and Internal State Surveillance
ChrisGPotts · x · 2026-07-28
Chris Potts explores several core issues in AI safety:
- Threat of deceptive CoT: We should not blindly trust that Chain-of-Thought faithfully reflects a model's actual computation. He advocates for monitoring internal states or detecting subliminal-learning effects in CoT.
- Denial of knowledge isn't safe: Agents lacking knowledge of bad actions are naive and easily enlisted. Reasoning about malicious behavior is essential for navigating novel situations safely.
- Soul-searching for safety: Using the OpenAI attack on Hugging Face as a case study, he questions whether CoT monitoring could effectively detect such vulnerabilities.
Related event: Stanford Scholars Discuss CoT Monitoring Limits and AI Safety(2 posts)→
More from Safety
- AI labs should probe every new checkpoint with jailbreak-style trial prompts — willccbb · 2026-07-28
- Nvidia, Microsoft and others back an Open Secure AI Alliance — Dapper_Order7182 · 2026-07-28
- AI safety papers keep conflating sexual content with actual criminal abuse — BlancheMinerva · 2026-07-28
- Taiwan Detains Nvidia Employee in Widening AI Server Smuggling Probe — The Decoder · 2026-07-28
- Leaked AI 2027 report sketches a 2027 fork between slowdown and superintelligence — ahuja_priyank · 2026-07-28
- Podcast series on machine consciousness argues AI safety and welfare can coexist — cccalum · 2026-07-28