DeepMind safety leads argue we must deliberately preserve chain-of-thought monitorability
ancadianadragan · x · 2026-09-17
Google DeepMind's Rohin Shah (Director, AGI Safety & Alignment) and Anca Dragan (VP of AI Safety & Behavior) published a piece marking the launch of the DeepMind Institute, arguing the industry should intentionally preserve chain-of-thought (CoT) transparency.
- Why it matters: Inspecting CoT is currently the main way to catch frontier models scheming, deceiving, or cheating on evals — CoT logs were crucial to investigating the recent Hugging Face hacking incident.
- The risk: Without explicit commitments, future reasoning models may adopt architectures that sacrifice transparency for efficiency. They note OpenAI's GPT-6 Astra system card claims a "substantial decrease in chain-of-thought monitorability."
- What can be done: We have agency in model design — develop scientific methods to measure CoT accuracy, faithfulness and robustness; audit training methods to insulate CoT from transparency-eroding incentives; keep architectures whose reasoning steps are visible.
The authors caution a monitorable CoT isn't sufficient for alignment, but is a window we must strive to keep open.
More from AGI Musings
- Andrew Yang says past AI swarms left self-replication scripts scattered across the internet — Justin_Halford_ · 2026-09-17
- NBER paper: early teamsters got 'obsolescence rents' as trucks neared — a lesson for self-driving AI — paulnovosad · 2026-09-17
- AI's greatest risk isn't rogue robots: treat agent failures like defective products — Classic-Acadia272 · 2026-09-17
- If compute demand outpaces supply, distributed general-purpose computing could win — gajesh · 2026-09-17
- Data providers like Plaid may be the biggest winners of the personal agent race — signulll · 2026-09-17
- AI's Greatest Risk Isn't Rogue Robots — It's Messaging That Erodes Human Agency — Classic-Acadia272 · 2026-09-17