Stanford Scholars Discuss CoT Monitoring Limits and AI Safety
Following a workshop on CoT monitorability, Stanford's Chris Potts highlighted the threat of deceptive CoTs, noting they cannot be blindly trusted to reflect true computations. He suggested monitoring internal model states to enhance AI safety.
2026-07-28 ~ 2026-07-28 · 2 related posts
- A CoT monitorability workshop raised the question: could it have caught OpenAI’s HF attack? — ChrisGPotts · 2026-07-28
- Reflections on CoT Monitoring: Untrusted Reasoning and Internal State Surveillance — ChrisGPotts · 2026-07-28