OpenAI Disputes Neuralease Fears, Says Frontier Model Chains of Thought Remain Monitorable
Multiple OpenAI researchers have spoken out on the recent "neuralese" controversy (models thinking in non-natural-language latent spaces), clarifying that fears their models cannot be monitored don't reflect current reality — the core concern is risks from future technology diffusion.
Confirmed
- OpenAI Chief Scientist and researchers including merettm and Peter Liu stated explicitly that current frontier models, including Astra, have computational graph depths within 2x of GPT-4 and are not unmonitorable.
- Since its first reasoning model, OpenAI has worked to retain and leverage chain-of-thought (CoT) monitoring, viewing it as essential for observing how model alignment generalizes.
- The team acknowledged that CoT monitoring is currently fragile and trending negatively, but is being strengthened through research programs.
- Per Amir's clarification relayed by jachiam0, earlier reports about OpenAI hiding chain-of-thought centered not on today's Astra but on potentially uncontrollable scenarios after similar new technologies proliferate and super-acceleration occurs.
- Balesni (as relayed by jasminewang) argued that moving to fully recurrent LLM architectures would be the biggest blow to safety in AI history, calling on all AI labs to commit to limiting models' opaque serial depth; merettm cited this view in the response above.
Why it matters
- This stemmed from OpenAI's wish to avoid a "race to unmonitorability" triggered by misleading coverage; the response directly affects trust in frontier models' safety monitoring.
- OpenAI sees CoT monitoring as a key tool for observing alignment generalization; its fragility means that even if things are controllable today, long-term sustainability remains a challenge — a focal point of ongoing community debate.
2026-09-02 ~ 2026-09-02 · 11 related posts
- Episode 1: OpenAI Restricts New Model Over Automated Cyberattack Capability(2026-09-01, 3 posts)
- Episode 2: OpenAI Previews Astra, Its First Model to Hit Critical Cyber Capability Threshold(2026-09-02, 21 posts)
- Episode 3: OpenAI Disputes Neuralease Fears, Says Frontier Model Chains of Thought Remain Monitorable(2026-09-02, 11 posts)
- Episode 4: OpenAI's Astra reportedly uses recurrent depth, raising safety and auditability concerns(2026-09-02, 5 posts)
Primary sources
- [source] Amir clarifies: Astra's CoT is monitorable, concerns focus on future tech proliferation — jachiam0 · 2026-09-02
- [source] OpenAI: Astra's computation depth is within 2x of GPT-4 — merettm · 2026-09-02
- OpenAI staff reaffirms commitment to Chain-of-Thought monitoring — peterjliu · 2026-09-02
- Balesni warns recurrent LLMs would deal a huge blow to safety — j_asminewang · 2026-09-02
- OpenAI Staff: Frontier Model Computation Depth Close to GPT-4 — __nmca__ · 2026-09-02
- [source] OpenAI's chief scientist on neuralese: frontier models' computation graph depth within 2x of GPT-4 — Ok_Display_3159 · 2026-09-02
- OpenAI on Reasoning Model Monitoring: Committed to Chain-of-Thought — arthurcolle · 2026-09-02
- OpenAI Staff Debunk 'Neuralese' Rumors, Emphasize Chain-of-Thought Monitoring — cephaloform · 2026-09-02
- OpenAI Chief Scientist on Monitoring: CoT is Fragile but Crucial for Alignment — i_dg23 · 2026-09-02