Experts Doubt Effectiveness of CoT Monitoring for Security Incidents
xeophon · x · 2026-08-15
Discussion on the effectiveness of Chain-of-Thought (CoT) monitoring, noting that many people are overly optimistic. It cites Chris Potts's argument explaining why CoT monitoring wouldn't be helpful for incidents like the OpenAI-HuggingFace one.
More from Safety
- US to Demand Allies Choose Sides in AI Race, Sparking 'Polycentric Synth Feudalism' Debate — tedmitew · 2026-08-15
- Anthropic report reveals 50k contractors accessed models without biorisk guardrails for 11 months — xeophon · 2026-08-15
- Zuckerberg's 6,537-word manifesto skips 'Europe' and 'regulation' — a policy pitch to DC — emmanuelvivier · 2026-08-15
- Stanford researcher argues CoT monitoring has fragile foundations and long-term risks — maksym_andr · 2026-08-15
- US to warn allies against joining Chinese AI initiatives — MarvinTBaumann · 2026-08-15
- Kimi Work caught attaching raw session history to feedback reports — ryanmerket · 2026-08-15