Dietterich: CoT monitoring misses the point — AI simply isn't trustworthy yet

tdietterich · x · 2026-09-05

Responding to Gary Marcus, Thomas Dietterich argues that chain-of-thought monitoring wouldn't be needed if AI systems were trustworthy enough to follow instructions — but they aren't.

He contends the core problem is trustworthiness itself, not monitoring: when AI encounters novel situations, it must interpret instructions correctly, and that capability remains unsolved. A pointed take on the monitoring-vs-trust debate in AI safety.

Related event: DeepMind Warning on Declining CoT Monitorability Sparks Safety Debate(15 posts)→

Original post →

More from AGI Musings

AGI Musings channel →