Back-of-Envelope: Frontier US Model Estimated at 8e26 FLOPs, 800B Active Params, 170T Training Tokens
teortaxesTex · x · 2026-09-15
TeortaxesTex extrapolates from Liang Wenfeng's May figure about the largest model known to be in development by Americans: assuming fp4 precision on Blackwells, 8.17e26 FLOPs, 800B active params, and a reasonable 4%-5% sparsity (16-20T total params), such a model would need roughly 170T training tokens. He adds that OpenAI can already access 100T+ tokens (citing glm-oss-related data). Speculative but source-anchored estimate of frontier model scale and training resources.
More from Infra
- Banning data centers to save the world? A 500-year history lesson says otherwise — thursdai_pod · 2026-09-15
- Forking Running Machines: Clone a Live Doom Session into 32 Realities in One Second — bigaiguy · 2026-09-15
- Running Qwen3.8-Flash-Next 125B MoE on a 12GB RTX 4070 at ~20 tok/s with MTP — carteakey · 2026-09-15
- Zuckerberg publicly asks Nvidia for a short-lead-time DGX Station for local AI — beffjezos · 2026-09-15
- BIS annual report picked apart: no codified H20 rule, Entity List stalled, loopholes open — ohlennart · 2026-09-15
- Apple's iOS 27 on-device AFM 3 has 20B params, activating only 1-4B — rxwei · 2026-09-15