OpenAI's Astra 'opaque reasoning' may severely impair AI safety oversight

jammastergirish · x · 2026-09-02

RyanGreenblatt reports that OpenAI's new AI, Astra, uses an 'opaque reasoning' architecture where reasoning occurs in activations rather than natural language. This is considered potentially the worst development for AI safety to date, as it could strongly impair oversight and monitoring. While current versions still rely on natural language CoT to some extent, there is concern that scaling up opaque reasoning to latent space would destroy the usefulness of monitoring.

Related event: OpenAI's Astra Reportedly Uses Recurrent Depth Architecture, Hiding Its Reasoning and Raising Safety Alarms(33 posts)→

Original post →

More from Safety

Safety channel →