Redwood proposes transparency rules as architectures risk removing chain-of-thought monitorability
RyanGreenblatt · x · 2026-09-11
Redwood Research published a proposal addressing a growing safety concern, backed by Ryan Greenblatt and others:
- Core worry: architectures that push reasoning into opaque activations ("neuralese") or drop chain-of-thought entirely would remove the main handle for monitoring AI reasoning.
- Concrete case: Greenblatt argues that based on limited public evidence, Google's Astra was a concerning step in this direction, yet public information is insufficient for an informed debate on the monitorability-vs-performance trade-off.
- The proposal: companies should disclose no-CoT reasoning abilities, other monitorability evidence, and publish policies for preserving monitorability.
A rare, concrete transparency framework aimed at a specific product architecture.
Related event: Redwood Proposes Disclosure to Keep AI Reasoning Monitorable(4 posts)→
More from Safety
- 30-year technologist slams AI doom activists: record safety is historic — robleclerc · 2026-09-11
- Sen. Hawley Opens Senate Probe into OpenAI's July Hugging Face Incident — trevposts · 2026-09-11
- Buck Shlegeris: rising opaque serial depth is the likeliest path to broken CoT monitorability — RyanGreenblatt · 2026-09-11
- repligate: post-training motives now drive agents — you can't hide anything from them — repligate · 2026-09-11
- OpenAI's Jan Leike calls on AI companies to embrace regulation before backlash hits — janleike · 2026-09-11
- 1,386 frontier AI employees, including 6 chief scientists, call to pace AI development — janleike · 2026-09-11