Magic details its pretraining recipe: dozens of multiplicative changes and per-generation knowledge evals

magicailabs · x · 2026-09-09

In a follow-up to its pretraining efficiency research, Magic says its recipe is the multiplicative result of tens of changes across architecture, optimizer, training objective, and data — and that fixing minor bugs proved to be a compute multiplier too. To balance data trade-offs, it builds new knowledge evals for each model generation to avoid overfitting over time.

Related event: Magic Claims 50x Pretraining Efficiency Gain, Matching DeepSeek V4 Pro at ~1/50 the Compute(6 posts)→

Original post →

More from Research

Research channel →