OpenAI's Astra 'opaque reasoning' may severely impair AI safety oversight
jammastergirish · x · 2026-09-02
RyanGreenblatt reports that OpenAI's new AI, Astra, uses an 'opaque reasoning' architecture where reasoning occurs in activations rather than natural language. This is considered potentially the worst development for AI safety to date, as it could strongly impair oversight and monitoring. While current versions still rely on natural language CoT to some extent, there is concern that scaling up opaque reasoning to latent space would destroy the usefulness of monitoring.
More from Safety
- Safin-1: Achieving Internal Safety via Memory-Native State Evolution — Shanghai-AI-Laboratory · 2026-09-02
- Gary Marcus warns OpenAI may cross AI safety redline — GaryMarcus · 2026-09-02
- GaryMarcus warns OpenAI reportedly sacrificing CoT monitorability for performance — AndyMasley · 2026-09-02
- Experts Worry Secret AI Communication Will Hinder Incident Investigations — sjgadler · 2026-09-02
- Ryan Greenblatt: Opaque Reasoning Architectures Are Extremely Bad for AI Safety — RyanGreenblatt · 2026-09-02
- VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models — _reachsumit · 2026-09-02