FULL STORY
Magic Claims DeepSeek V4 Pro Pretraining Match at 1/50 the Compute
Magic released a pretraining efficiency update claiming to match DeepSeek V4 Pro pretraining with roughly 1/50 the FLOPs, then outlined its next steps: long-context RL for test-time learning and alignment before model release.
2026-09-09 ~ 2026-09-09 · 2 episodes · 13 posts
Episode 1 · Magic Claims to Match DeepSeek V4 Pro Pretraining with 1/50 the Compute (2026-09-09, 11 posts)
On September 9, AI startup Magic published a pretraining efficiency research update claiming that algorithmic efficiency can close the compute gap with giants: a new recipe matches DeepSeek V4 Pro Base pretraining with roughly 1/50 of the FLOPs, or about $500K of GB200 compute (roughly half of GPT-3's pretraining compute). The team says it lacks 100K chips and that efficiency is its only path, challenging the assumption that frontier pretraining is a big-lab game. All figures remain Magic's own claims with no third-party replication.
Confirmed
- Core claim: matching DeepSeek V4 Pro Base with 1/50 FLOPs and $500K of GB200 compute; per DrSingularity's summary, the recipe is over 10x more compute-efficient than leading open-weight training methods, with future targets of trillion-parameter scale and automated AI R&D.
- Cost framing: reaching equivalent capability with DeepSeek V4 Pro's recipe would cost over $100M; Magic says $4M (10x scale) surpassed all publicly available base models.
- Methodology: the recipe is the multiplicative result of dozens of changes across architecture, optimizer, training objectives, and data; fixing small bugs is itself a compute multiplier.
- Evaluation: Magic targets strongest SWE and autonomous AI R&D models, balancing data tradeoffs with standard out-of-sample suites plus knowledge evals redone each generation to avoid overfitting benchmarks.
Why it matters
- If verified, small teams could compete on frontier base models via algorithmic efficiency rather than massive compute.
- The multiplicative-changes framing elevates engineering detail to strategic importance, valuable for resource-constrained teams.
- Regenerating knowledge evals each generation offers a reusable practice for evaluation credibility.
- All data are unverified single-source claims; treat official blog follow-ups as the reference.
- Magic claims 50x pretraining efficiency: matches DeepSeek V4 Pro for ~$0.5M — magicailabs · 2026-09-09
- MagicaLabs says $4M pretraining beat all public base models, 50x more compute-efficient than DeepSeek — magicailabs · 2026-09-09
- MagicaLabs details its pretraining recipe: tens of multiplicative changes, bug fixes as compute multipliers — magicailabs · 2026-09-09
- Magic details its pretraining recipe: dozens of multiplicative changes and per-generation knowledge evals — magicailabs · 2026-09-09
- MagicAI claims pretraining recipe matches frontier model with 50x less compute, ~$0.5M — EricSteinb · 2026-09-09
- Magic claims new recipe matches DeepSeek V4 Pro pretraining with 50x less compute (~$0.5M) — daniel_mac8 · 2026-09-09
- MagicAILabs claims new recipe matches DeepSeek V4 Pro pretraining with 50x less compute, ~$0.5M — AccBalanced · 2026-09-09
- Magic claims >10x more efficient pretraining, matches DeepSeek V4 Pro with 50x fewer FLOPs — Dr_Singularity · 2026-09-09
- Magic details >10x compute-efficient pretraining, eyes trillion-parameter models — Dr_Singularity · 2026-09-09
- Magenta claims it matched DeepSeek V4 Pro pretraining with 50x less compute, ~$0.5M — generativist · 2026-09-09
- Magic rebuilds knowledge evals per model generation to avoid eval overfitting — seanmcdonaldxyz · 2026-09-09
Episode 2 · Magic Outlines Next Steps and Calls Itself the Smallest Trillion-Parameter Team (2026-09-09, 2 posts)
After its pretraining research update, Magic plans to extend RL with long context for test-time learning, align models using latent knowledge of their own intentions, then release models, while calling itself possibly the smallest team training trillion-parameter models and announcing hiring.
- Magic's roadmap: long-context RL, latent-knowledge alignment, then a model release — magicailabs · 2026-09-09
- Magic, likely the world's smallest team training trillion-parameter models, is hiring — magicailabs · 2026-09-09