MagicaLabs says $4M pretraining beat all public base models, 50x more compute-efficient than DeepSeek
magicailabs · x · 2026-09-09
MagicaLabs claims its new pretraining recipe, scaled 10x for roughly $4M, exceeds all publicly available base models — a capability that would cost over $100M under DeepSeek V4 Pro's recipe.
- The bet is algorithmic efficiency over chip count: the recipe matches DeepSeek V4 Pro's pretraining with 50x less compute, roughly half the FLOPs of GPT-3, or $0.5M on GB200.
- The gains come from tens of multiplicative changes across architecture, optimizer, training objective, and data; the team found fixing minor bugs is itself a compute multiplier.
- A short math RL run from base was used to sanity-check post-RL performance, and the team says it won't stop scaling there.
- These are self-reported claims with no third-party replication yet.
More from Infra
- Fab2 raises $500M Series A at $3.7B valuation to scale chip fab business — Sethwinterroth · 2026-09-09
- Zach Dell on Base Power: meeting AI's energy demands, batteries and vertical integration — espricewright · 2026-09-09
- Smartphone brands raise prices in India as DRAM shortage seen lasting until late 2027 — saibharadwaj · 2026-09-09
- Inception ships Mercury 2.5: most capable diffusion LLM at 1,107 tokens/sec, 260K context — StefanoErmon · 2026-09-09
- Intel shows a 27x speedup in matrix multiplication by just swapping two loops, both O(n³) — jedisct1 · 2026-09-09
- AWS benchmarks G7 Blackwell vs G5/G6 for 30B MoE inference on SageMaker — AWS ML Blog · 2026-09-09