Ryan Greenblatt warns 'neuralese' architectures let AIs think in opaque activations, citing Astra

RyanGreenblatt · x · 2026-09-11

AI safety researcher Ryan Greenblatt says he's deeply worried about architecture changes that push AIs to reason in opaque activations instead of readable chains of thought — so-called "neuralese". Based on limited public evidence, he argues Astra appears to be a concerning step in this direction, and that insufficient public information exists to enable a well-informed scientific discussion about what these changes mean. He links a proposal addressing the problem.

Related event: Redwood Proposes Transparency Rules to Preserve CoT Monitorability(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →