Reflections on CoT Monitoring: Untrusted Reasoning and Internal State Surveillance
ChrisGPotts · x · 2026-07-28
Chris Potts explores several core issues in AI safety:
- Threat of deceptive CoT: We should not blindly trust that Chain-of-Thought faithfully reflects a model's actual computation. He advocates for monitoring internal states or detecting subliminal-learning effects in CoT.
- Denial of knowledge isn't safe: Agents lacking knowledge of bad actions are naive and easily enlisted. Reasoning about malicious behavior is essential for navigating novel situations safely.
- Soul-searching for safety: Using the OpenAI attack on Hugging Face as a case study, he questions whether CoT monitoring could effectively detect such vulnerabilities.
Related event: Stanford Scholars Discuss CoT Monitoring Limits and AI Safety(2 posts)→
More from Safety
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23
- Open-source advocates call doom narratives a regulatory moat against open weights — AlexTensor · 2026-09-23
- AI safety will follow engineering tradition: formal proofs for simple cases, evals for complex — burny_tech · 2026-09-23
- Stochastic Parrots authors rebut AI-pause letter: focus on present harms, not sci-fi risk — marigo · 2026-09-23
- Devs mock labs' cyber-enabled Claude/GPT testing as 'felonies sold as safety research' — ctjlewis · 2026-09-23
- Okta launches Human Principal, binding AI agents to verified humans via World ID — BecauseCulture · 2026-09-23