One cluster alone could train 68 GPT-6-scale models by 2029, and FP4 could double that

scaling01 · x · 2026-09-04

Quoting his own compute-scale tweet, scaling01 calculates that a single large cluster among dozens could support training roughly 68 GPT-6-scale runs by 2029, remarking "we are still so early." He adds he hasn't even factored in FP4 precision, which could make the number 2x-4x higher — an exaggerated but illustrative take on the sheer scale of AI compute buildout and headroom from precision gains.

Original post →

More from Infra

Infra channel →