Magenta claims it matched DeepSeek V4 Pro pretraining with 50x less compute, ~$0.5M

generativist · x · 2026-09-09

Magenta AI (@magicailabs) claims that with no 100k-chip cluster, algorithmic efficiency is the only path to frontier pretraining. Its new recipe reportedly matches DeepSeek V4 Pro's pretraining using 50x less compute — roughly half the FLOPs of GPT-3, or about $0.5M on GB200. Widely shared as "big if true," but unverified; details hinge on reproducibility.

Related event: Magic Claims to Match DeepSeek V4 Pro Pretraining with 1/50 the Compute(11 posts)→

Original post →

More from Infra

Infra channel →