OpenAI's Astra reported to use 'opaque reasoning' architecture, harming safety oversight
thlarsen · x · 2026-09-02
Ryan Greenblatt comments on reports that OpenAI's new model, Astra, uses an 'opaque reasoning' architecture where reasoning occurs in activations rather than natural language. He labels this potentially the single worst development for AI safety to date.
Key Points:
- Oversight Risks: This architecture could significantly impair supervision and monitoring, as reasoning processes are no longer fully visible in natural language Chain-of-Thought (CoT).
- Current Status: The article suggests recurrent depth is limited, meaning the model still relies on natural language CoT, though to a reduced extent, making text-based monitoring much less useful.
- Future Concerns: A natural progression scaling up opaque reasoning could lead to models reasoning entirely in latent space, likely destroying the usefulness of natural language reasoning for oversight.
More from AGI Musings
- Joseph Jacks urges more long-term planning for 2040s and 2050s — JosephJacks_ · 2026-09-02
- AI Reading Every Paper on arXiv Could Be a Bridge to AGI — imjustnewatai · 2026-09-02
- Boaz Barak: Centralized ASI increases misaligned singleton risk — aidan_mclau · 2026-09-02
- Model Scaling Trend: 2T Parameters Becoming New Norm as KV Cache Shrinks 10x YoY — zephyr_z9 · 2026-09-02
- Former Tech PR Turned AI Bear Warns of Economic Bubble Burst — whurley · 2026-09-02
- Study: ChatGPT caused 21-50% drop in writing variance across the web — maier_ak · 2026-09-02