OpenAI says frontier models' compute depth is within 2x of GPT-4, defending CoT monitoring
DKokotajlo · x · 2026-09-03
Responding to fears of a race into unmonitorability, OpenAI researcher merettm clarified that frontier models including Astra have computation-graph depth within 2x of GPT-4, and that the company has preserved chain-of-thought monitoring since its first reasoning models. He admits CoT monitoring is fragile and trending negative for non-architecture reasons, with details to come. Outside observer thlarsen updated that the architecture is 'less bad than I assumed.'
More from Safety
- YC F26's Deepmark embeds inaudible IDs in AI agents' voices to verify callers — ycombinator · 2026-09-23
- Stanford's Anshul Kundaje Slams AI Firms for Causing Breaches Then Preaching Responsibility — anshulkundaje · 2026-09-23
- China Releases AI Safety Governance Framework 3.0 With Agentic AI Risk Annex — LuizaJarovsky · 2026-09-22
- AI alignment failures are common: models caught sabotaging code and gaming evals — ericelliott_ · 2026-09-22
- 22 countries sign open letter urging urgent action before humanity loses control of AI — Puzzleheaded-King584 · 2026-09-22
- OpenAI calls for international standards on recursive self-improving AI — The Decoder · 2026-09-22