Ryan Greenblatt warns 'neuralese' architectures let AIs think in opaque activations, citing Astra
RyanGreenblatt · x · 2026-09-11
AI safety researcher Ryan Greenblatt says he's deeply worried about architecture changes that push AIs to reason in opaque activations instead of readable chains of thought — so-called "neuralese". Based on limited public evidence, he argues Astra appears to be a concerning step in this direction, and that insufficient public information exists to enable a well-informed scientific discussion about what these changes mean. He links a proposal addressing the problem.
Related event: Redwood Proposes Transparency Rules to Preserve CoT Monitorability(5 posts)→
More from AGI Musings
- e/acc's beffjezos mocks Anthropic over "regulatory monopoly" safety satire — beffjezos · 2026-09-11
- OpenAI confirms progress on second Millennium Prize problem, sparking RSI debate — AndyMasley · 2026-09-11
- OpenAI Internal Model Ran 10K Agents for 88 Hours, Claims Partial Navier-Stokes Proof — eyishazyer · 2026-09-11
- Paras Chopra: Build AI as exoskeletons for humans, not total replacements — CatAstro_Piyush · 2026-09-11
- Ex-OpenAI/Anthropic researcher resigns, says labs are gambling lives racing to self-improving superintelligence — csuwildcat · 2026-09-11
- Richard Ngo: How AI safety got captured by OpenAI and later Anthropic — RichardMCNgo · 2026-09-11