Redwood Proposes CoT Monitorability Commitments Amid 'Neuralese' Speedrun Rumors
RyanGreenblatt · x · 2026-09-11
Rumors that labs were 'speedrunning neuralese' — not true yet — have made concrete CoT monitorability commitments more urgent, says Anthropic's Ryan Greenblatt, sharing Redwood Research's new proposal.
- Redwood warns some architectures could weaken CoT monitorability or remove chain-of-thought entirely, as incentives naturally slide toward opaque architectures
- The proposal calls on companies to be transparent about no-CoT reasoning abilities, disclose other monitorability evidence, and adopt policies preserving monitorability
- Greenblatt hopes OpenAI, Anthropic, and others will rally around it
Related event: Redwood Proposes Transparency Rules to Preserve CoT Monitorability(5 posts)→
More from Safety
- Shared AI Chat Links Are Not as Private as You Think — RummanSid1990 · 2026-09-11
- Microsoft Patches Record 974 Vulnerabilities, Mostly Found by AI — Distinct-Question-16 · 2026-09-11
- Anthropic's September 2026 Threat Intelligence Report on AI Misuse — Cubewood · 2026-09-11
- Ben Bajarin: agentic AI in cyber defense is the next frontier, but authority limits remain the challenge — BenBajarin · 2026-09-11
- Method and Palantir launch Cardinal Program: free autonomous AI security assessments for critical infrastructure — adilmajid · 2026-09-11
- Gary Marcus: Doom talk distracts from frontier labs' incompetent security engineering — GaryMarcus · 2026-09-11