Magic claims 50x pretraining efficiency: matches DeepSeek V4 Pro for ~$0.5M

magicailabs · x · 2026-09-09

Magic published a research update claiming its pretraining recipe is >10x more compute-efficient than leading open-weight base models: it matches DeepSeek V4 Pro Base with 50x fewer FLOPs (half of GPT-3's pretraining compute, $0.5M on GB200), and scaling 10x more ($4M) meaningfully beats all public open base models on perplexity. Its scaling laws imply DeepSeek's recipe would cost >$100M to reach the same capability. Evaluated bits-per-byte loss across DeepSeek, Kimi and NVIDIA base models on GB200/GB300 with vLLM and SGLang.

Related event: Magic Claims 50x Pretraining Efficiency Gain, Matching DeepSeek V4 Pro at ~1/50 the Compute(6 posts)→

Original post →

More from Infra

Infra channel →