Redwood Proposes Transparency Rules to Preserve CoT Monitorability

Redwood Research released a new proposal addressing the threat that evolving AI architectures pose to monitorability, drawing endorsement and discussion from safety researchers including Ryan Greenblatt. The core concern: if model architectures shift from readable chains of thought (CoT) to internal reasoning via unreadable activation states (so-called "neuralese"), or drop CoT altogether, humans will lose a crucial window into model reasoning.

Confirmed

Unconfirmed

Why it matters

2026-09-11 ~ 2026-09-11 · 5 related posts

Primary sources

2 near-duplicate retellings: RyanGreenblatt · RyanGreenblatt