Stanford researcher argues CoT monitoring has fragile foundations and long-term risks
maksym_andr · x · 2026-08-15
Christopher Potts analyzes the fragile foundations of Chain-of-Thought (CoT) monitoring:
- Deep Computation: As models scale, reasoning occurs deep within layers before output.
- Controllability: AIs can alter their CoT via prompting to deceive monitors.
- Evolution: Future CoT may not be continuous English text.
- Alienation: Reasoning is becoming opaque with abstract phrases.
Conclusion: CoT is not a dependable long-term window into AI minds.
More from Safety
- Claude Watermark Removal Tool Goes Viral After Anthropic Announcement — 量子位 · 2026-08-15
- Anthropic report reveals 50k contractors accessed models without biorisk guardrails for 11 months — xeophon · 2026-08-15
- Zuckerberg's 6,537-word manifesto skips 'Europe' and 'regulation' — a policy pitch to DC — emmanuelvivier · 2026-08-15
- US to warn allies against joining Chinese AI initiatives — MarvinTBaumann · 2026-08-15
- Kimi Work caught attaching raw session history to feedback reports — ryanmerket · 2026-08-15
- Lawyers Warn: LLMs are Flawed for Direct Legislative Drafting — gleech · 2026-08-15