OpenAI's Astra reported to use 'opaque reasoning' architecture, harming safety oversight

thlarsen · x · 2026-09-02

Ryan Greenblatt comments on reports that OpenAI's new model, Astra, uses an 'opaque reasoning' architecture where reasoning occurs in activations rather than natural language. He labels this potentially the single worst development for AI safety to date.

Key Points:

Related event: OpenAI's Astra Reportedly Uses Recurrent Depth Reasoning, Hiding Chain of Thought and Raising Safety Concerns(49 posts)→

Original post →

More from AGI Musings

AGI Musings channel →