GPT-6 could use 2-5x pre-training compute for RL, speculates scaling01

scaling01 · x · 2026-09-07

AI commentator scaling01 argues that continuous 100-day pre-training runs are obsolete since they waste algorithmic progress, and floats a base-case guess for GPT-6 ("Astra"): on top of pre-training, roughly 2x-5x the pre-training compute spent on RL. Pure speculation, not confirmed.

Related event: KOL Tests Astra, Sees No AGI; Speculates GPT-6 Uses Up to 5x Compute(3 posts)→

Original post →

More from Models

Models channel →