Researchers warn against "neuralese" architectures that trade AI monitorability for performance

RyanGreenblatt · x · 2026-09-11

Redwood Research's Ryan Greenblatt voiced concern about AI architectures shifting from readable chain-of-thought to opaque "neuralese" activations. Based on limited public evidence, he sees Google's Astra as a concerning step in this direction: the monitorability-vs-performance trade-offs in its architecture and training lack sufficient public evidence for informed scientific discussion, and it's unclear how AI companies will make such trade-offs going forward. He calls on companies to release evidence enabling a reasonably informed public conversation. OpenAI researcher balesni amplified the thread, saying he thinks AI is >10% likely to kill all humans and this proposal is among the industry's top risk-reduction steps.

Related event: Redwood Proposes Transparency Rules to Preserve CoT Monitorability(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →