Neuralese explained: why OpenAI's Astra architecture has safety researchers alarmed

ShakeelHashim · x · 2026-09-04

Following a report that OpenAI's new Astra model uses an architecture that could make its reasoning harder to monitor, this explainer covers why chain-of-thought matters for both capability and safety, what 'neuralese' means, and why researcher Ryan Greenblatt called the news possibly 'the single worst development for AI security/safety to date.'

Related event: Researchers Debate Whether Chain-of-Thought Monitoring Can Keep AI Safe(38 posts)→

Original post →

More from Models

Models channel →