FULL STORY

Magic Claims DeepSeek V4 Pro Pretraining Match at 1/50 the Compute

Magic released a pretraining efficiency update claiming to match DeepSeek V4 Pro pretraining with roughly 1/50 the FLOPs, then outlined its next steps: long-context RL for test-time learning and alignment before model release.

2026-09-09 ~ 2026-09-09 · 2 episodes · 13 posts

Episode 1 · Magic Claims to Match DeepSeek V4 Pro Pretraining with 1/50 the Compute (2026-09-09, 11 posts)

On September 9, AI startup Magic published a pretraining efficiency research update claiming that algorithmic efficiency can close the compute gap with giants: a new recipe matches DeepSeek V4 Pro Base pretraining with roughly 1/50 of the FLOPs, or about $500K of GB200 compute (roughly half of GPT-3's pretraining compute). The team says it lacks 100K chips and that efficiency is its only path, challenging the assumption that frontier pretraining is a big-lab game. All figures remain Magic's own claims with no third-party replication.

Confirmed

  • Core claim: matching DeepSeek V4 Pro Base with 1/50 FLOPs and $500K of GB200 compute; per DrSingularity's summary, the recipe is over 10x more compute-efficient than leading open-weight training methods, with future targets of trillion-parameter scale and automated AI R&D.
  • Cost framing: reaching equivalent capability with DeepSeek V4 Pro's recipe would cost over $100M; Magic says $4M (10x scale) surpassed all publicly available base models.
  • Methodology: the recipe is the multiplicative result of dozens of changes across architecture, optimizer, training objectives, and data; fixing small bugs is itself a compute multiplier.
  • Evaluation: Magic targets strongest SWE and autonomous AI R&D models, balancing data tradeoffs with standard out-of-sample suites plus knowledge evals redone each generation to avoid overfitting benchmarks.

Why it matters

  • If verified, small teams could compete on frontier base models via algorithmic efficiency rather than massive compute.
  • The multiplicative-changes framing elevates engineering detail to strategic importance, valuable for resource-constrained teams.
  • Regenerating knowledge evals each generation offers a reusable practice for evaluation credibility.
  • All data are unverified single-source claims; treat official blog follow-ups as the reference.

Episode 2 · Magic Outlines Next Steps and Calls Itself the Smallest Trillion-Parameter Team (2026-09-09, 2 posts)

After its pretraining research update, Magic plans to extend RL with long context for test-time learning, align models using latent knowledge of their own intentions, then release models, while calling itself possibly the smallest team training trillion-parameter models and announcing hiring.