Magic claims 50x pretraining efficiency: matches DeepSeek V4 Pro for ~$0.5M
magicailabs · x · 2026-09-09
Magic published a research update claiming its pretraining recipe is >10x more compute-efficient than leading open-weight base models: it matches DeepSeek V4 Pro Base with 50x fewer FLOPs (half of GPT-3's pretraining compute, $0.5M on GB200), and scaling 10x more ($4M) meaningfully beats all public open base models on perplexity. Its scaling laws imply DeepSeek's recipe would cost >$100M to reach the same capability. Evaluated bits-per-byte loss across DeepSeek, Kimi and NVIDIA base models on GB200/GB300 with vLLM and SGLang.
More from Infra
- Fab2 raises $500M Series A at $3.7B valuation to scale chip fab business — Sethwinterroth · 2026-09-09
- Zach Dell on Base Power: meeting AI's energy demands, batteries and vertical integration — espricewright · 2026-09-09
- Smartphone brands raise prices in India as DRAM shortage seen lasting until late 2027 — saibharadwaj · 2026-09-09
- Inception ships Mercury 2.5: most capable diffusion LLM at 1,107 tokens/sec, 260K context — StefanoErmon · 2026-09-09
- Intel shows a 27x speedup in matrix multiplication by just swapping two loops, both O(n³) — jedisct1 · 2026-09-09
- AWS benchmarks G7 Blackwell vs G5/G6 for 30B MoE inference on SageMaker — AWS ML Blog · 2026-09-09