Back-of-Envelope Estimate Puts DeepSeek V4.1 Pretraining at ~5e24 FLOPs
teortaxesTex · x · 2026-09-10
Citing Astra's estimate, a poster puts DeepSeek V4.1's pretraining compute at 5e24 FLOPs — slightly above V3, or about 3.5M H100-hours, doable in 2 weeks on 4K B300s. With 16K B300s, they muse, one could train a 256B-active/10T-token model, or a 64B-active model at 3-4T tokens in 2 months. Unofficial back-of-envelope speculation.
Related event: Estimate: V4.1 Pretraining Took ~5e24 FLOPs(2 posts)→
More from Infra
- Positron AI raises $230M Series B at over $1B valuation with Arm backing — seanmcdonaldxyz · 2026-09-11
- Cerebras Fast Inference Flips Agent Workflows: Fewer Parallel Agents, Same Output — MatthewBerman · 2026-09-11
- Baseten acquires Blaxel to build integrated cloud infrastructure for AI agents — baseten · 2026-09-11
- Hyperscalers could factor RSA-1024 for about $30M per number, analysis claims — rickasaurus · 2026-09-11
- Qualcomm's Next Hexagon NPU: 50% More Shared Memory, 30B MoE Models on a Phone — ryanshrout · 2026-09-11
- Vercel Cut CDN P99 Metadata Lookup Latency by 91% Across 80M Route Decisions/sec — cramforce · 2026-09-11