Magic claims >10x more efficient pretraining, matches DeepSeek V4 Pro with 50x fewer FLOPs
Dr_Singularity · x · 2026-09-09
Startup Magic says its new pretraining recipe is over 10x more compute-efficient than leading open-weight models.
- Matches DeepSeek V4 Pro Base with 50x fewer FLOPs and $0.5M of GB200 compute
- Scaling to $4M produced a base model that beat all public open models on perplexity evals; estimated cost under DeepSeek's recipe: >$100M
- Magic is scaling toward trillion-parameter models and autonomous AI R&D
More from Infra
- Baseten tops Coval's voice AI benchmark: STT ~5x faster than OpenAI with lowest WER — baseten · 2026-09-09
- 2×4090 llama.cpp concurrency: soft cap of 5 agents at 64k context, hard cap 9 — three weeks of data — Iamisseibelial · 2026-09-09
- exe.dev deep dive: ssh to a persistent Linux VM in half a second, priced like a folder — davidcrawshaw · 2026-09-09
- GPT-6 Astra lands on Amazon Bedrock with 1M-token context and first Critical cyber rating — AWS ML Blog · 2026-09-09
- DeepSeek v4 and GLM Now Run Faster Than vLLM and SGLang — jedisct1 · 2026-09-09
- Estha Turns One Mac Into a Shared Local AI Server for a Whole Team — HaktanSuren · 2026-09-09