Redwood Proposes Transparency Rules to Preserve CoT Monitorability
Redwood Research released a new proposal addressing the threat that evolving AI architectures pose to monitorability, drawing endorsement and discussion from safety researchers including Ryan Greenblatt. The core concern: if model architectures shift from readable chains of thought (CoT) to internal reasoning via unreadable activation states (so-called "neuralese"), or drop CoT altogether, humans will lose a crucial window into model reasoning.
Confirmed
- Redwood Research published a proposal warning that certain architectures could weaken or entirely remove the monitorability of CoT.
- The proposal recommends that companies transparently disclose three things: non-CoT reasoning capabilities, other evidence of monitorability, and policies for maintaining monitorability.
- Greenblatt said that based on limited public information, he views Google's Astra as a concerning relevant case.
Unconfirmed
- Earlier rumors claimed labs were racing to develop "neuralese" approaches; per shared posts, this is reportedly not true and remains unverified.
Why it matters
- CoT monitorability is considered one of the key current tools for humans to supervise powerful models. If architectures move toward unreadable internal reasoning, the toolbox for safety evaluation and alignment verification will shrink considerably. The proposal aims to preserve this monitoring baseline amid the architecture race through transparent disclosure obligations.
2026-09-11 ~ 2026-09-11 · 5 related posts
Primary sources
- Redwood proposes transparency rules as architectures risk removing chain-of-thought monitorability — RyanGreenblatt ·
- Ryan Greenblatt warns 'neuralese' architectures let AIs think in opaque activations, citing Astra — RyanGreenblatt ·
- Redwood proposal: companies should disclose no-CoT reasoning to preserve AI monitorability — RyanGreenblatt ·
- [source] Redwood proposes transparency rules as architectures risk removing chain-of-thought monitorability — RyanGreenblatt · 2026-09-11
- [source] Ryan Greenblatt warns 'neuralese' architectures let AIs think in opaque activations, citing Astra — RyanGreenblatt · 2026-09-11
- Redwood Proposes CoT Monitorability Commitments Amid 'Neuralese' Speedrun Rumors — RyanGreenblatt · 2026-09-11
2 near-duplicate retellings: RyanGreenblatt · RyanGreenblatt