OpenAI researcher: frontier models incl. Astra stay within 2x of GPT-4 compute depth

deanwball · x · 2026-09-02

DeepMind researcher Tomasz Korbak argues that the day a frontier lab trains a frontier-scale recurrent (or otherwise unmonitorable) LLM would be among the darkest of the current AI era, urging labs to coordinate on a commitment to never do so.

OpenAI's merettm pushed back on "confused reporting" fueling a race into unmonitorability: the computation-graph depth of today's frontier models, including Astra, is within a factor of two of GPT-4. He says OpenAI has preserved and leveraged chain-of-thought monitoring since its first reasoning models, as it reveals how alignment generalizes beyond the training distribution — but he admits the technique is fragile and trending negative for non-architectural reasons he'll detail soon, and strengthening it is a core research goal.

Related event: OpenAI researchers push back on neuralese fears, saying frontier models remain monitorable(20 posts)→

Original post →

More from AGI Musings

AGI Musings channel →