OpenAI says frontier models' compute depth is within 2x of GPT-4, defending CoT monitoring

DKokotajlo · x · 2026-09-03

Responding to fears of a race into unmonitorability, OpenAI researcher merettm clarified that frontier models including Astra have computation-graph depth within 2x of GPT-4, and that the company has preserved chain-of-thought monitoring since its first reasoning models. He admits CoT monitoring is fragile and trending negative for non-architecture reasons, with details to come. Outside observer thlarsen updated that the architecture is 'less bad than I assumed.'

Related event: OpenAI researchers push back on neuralese fears, saying frontier models remain monitorable(20 posts)→

Original post →

More from Safety

Safety channel →