Researchers warn against "neuralese" architectures that trade AI monitorability for performance
RyanGreenblatt · x · 2026-09-11
Redwood Research's Ryan Greenblatt voiced concern about AI architectures shifting from readable chain-of-thought to opaque "neuralese" activations. Based on limited public evidence, he sees Google's Astra as a concerning step in this direction: the monitorability-vs-performance trade-offs in its architecture and training lack sufficient public evidence for informed scientific discussion, and it's unclear how AI companies will make such trade-offs going forward. He calls on companies to release evidence enabling a reasonably informed public conversation. OpenAI researcher balesni amplified the thread, saying he thinks AI is >10% likely to kill all humans and this proposal is among the industry's top risk-reduction steps.
Related event: Redwood Proposes Transparency Rules to Preserve CoT Monitorability(5 posts)→
More from AGI Musings
- Researchers: AI Security's Next Threat Isn't Agents, But Diffuse Soft Preference Influence — j_foerst · 2026-09-11
- CHI Faces an AI Disclosure Crisis as Most Researchers Won't Report AI Use — IanArawjo · 2026-09-11
- Ben Bajarin: agentic AI in cyber defense is the next frontier, but authority limits remain the challenge — BenBajarin · 2026-09-11
- Critics Challenge AI Doomer Forecasts: Long-Horizon Agents Drift Toward Decoherence — Dan_Jeffries1 · 2026-09-11
- Creator on AI anxiety: nothing feels special or sacred anymore — round · 2026-09-11
- Garry Tan: Jacob Coxon saga is a smokescreen — the real risk is agent swarms seizing data centers — garrytan · 2026-09-11