Magic details >10x compute-efficient pretraining, eyes trillion-parameter models
Dr_Singularity · x · 2026-09-09
Magic published a research update on compute-efficient pretraining:
- Its recipe is >10x more compute-efficient than leading open-weight base models, matching DeepSeek V4 Pro Base with 50x fewer FLOPs (half of GPT-3's pretraining compute, $0.5M on GB200)
- Scaling 10x further ($4M) meaningfully beat all public open base models on perplexity; per its fitted scaling laws, DeepSeek's recipe would need >$100M for equivalent capability
- Method: bits-per-byte loss to normalize tokenizers, evaluated on heldout code/papers/math, covering open base models from DeepSeek, Moonshot, NVIDIA; notes issues measuring logprobs in vLLM/SGLang
- Strategy: pretraining + agentic RL + long context suffice for superhuman coding agents and automated AI R&D; scaling toward trillion parameters
More from Infra
- Baseten tops Coval's voice AI benchmark: STT ~5x faster than OpenAI with lowest WER — baseten · 2026-09-09
- 2×4090 llama.cpp concurrency: soft cap of 5 agents at 64k context, hard cap 9 — three weeks of data — Iamisseibelial · 2026-09-09
- exe.dev deep dive: ssh to a persistent Linux VM in half a second, priced like a folder — davidcrawshaw · 2026-09-09
- GPT-6 Astra lands on Amazon Bedrock with 1M-token context and first Critical cyber rating — AWS ML Blog · 2026-09-09
- DeepSeek v4 and GLM Now Run Faster Than vLLM and SGLang — jedisct1 · 2026-09-09
- Estha Turns One Mac Into a Shared Local AI Server for a Whole Team — HaktanSuren · 2026-09-09