OpenAI's new reasoning method may obscure CoT, sparking safety concerns
AaronBergman18 · x · 2026-09-02
Researchers identified that OpenAI's Astra AI uses a new reasoning approach called "recurrent depth." While this can improve performance and reduce costs, it obscures the model's thinking process, making monitoring difficult. This raises concerns as CoT monitoring is a key part of OpenAI's safety plan, appearing inconsistent with previous statements.
More from Safety
- Ilya Sutskever posts on security against rogue AI models — borowcy · 2026-09-02
- Model looping monitorability hinges on effective depth, not binary nature — teortaxesTex · 2026-09-02
- OpenAI's chief scientist on neuralese: frontier models' computation graph depth within 2x of GPT-4 — Ok_Display_3159 · 2026-09-02
- Call for collective standards on dangerous AI training — sjgadler · 2026-09-02
- CoT monitoring may fail: misaligned AI gets harder to detect, outpacing AI 2027 — AaronBergman18 · 2026-09-02
- Hacking SQL Server AI Assistant: From SELECT to SYSADMIN — wunderwuzzi23 · 2026-09-02