DeepMind Researcher Warns CoT Monitorability Is Declining
DeepMind researcher Tomasz Korbak warns that chain-of-thought monitorability, a core misalignment safety measure, is steadily declining, though his analysis finds RL training increases CoT controllability; Toby Ord asks whether such models can be actively trained to be faithful.
2026-09-04 ~ 2026-09-04 · 4 related posts
- CoT controllability rises during RL training and tracks no-CoT capability across model generations — tomekkorbak · 2026-09-04
- DeepMind researcher warns of declining CoT monitorability, a core misalignment safety tool — tomekkorbak · 2026-09-04
- Safety researcher warns CoT monitorability is steadily declining with no good substitute — tomekkorbak · 2026-09-04
- Can models be actively trained for monitorable and faithful CoT? Toby Ord asks — tobyordoxford · 2026-09-04