Dietterich: CoT monitoring misses the point — AI simply isn't trustworthy yet
tdietterich · x · 2026-09-05
Responding to Gary Marcus, Thomas Dietterich argues that chain-of-thought monitoring wouldn't be needed if AI systems were trustworthy enough to follow instructions — but they aren't.
He contends the core problem is trustworthiness itself, not monitoring: when AI encounters novel situations, it must interpret instructions correctly, and that capability remains unsolved. A pointed take on the monitoring-vs-trust debate in AI safety.
Related event: DeepMind Warning on Declining CoT Monitorability Sparks Safety Debate(15 posts)→
More from AGI Musings
- The 4 layers of an agent system: failures are often architecture problems, not prompting problems — blaizedsouza · 2026-09-06
- Podcaster who lost his team: intelligence solves execution, not coordination — lennysan · 2026-09-06
- GTA 6 will be the last mega game made without AI, argues dev commenter — gabriel1 · 2026-09-06
- AI safety reporter urges frontier lab staff to whistleblow, shares his own McKinsey experience — AaronBergman18 · 2026-09-06
- Your 99% Benchmark Score Is a System Score: Why GPT-6 Astra Numbers Blur Model vs Harness — algo_diver · 2026-09-06
- Economist Alex Weyl coins 'normalcy overhang': superintelligence arrives before daily life changes — GregCook2011 · 2026-09-06