A rough compute check puts a training run at 3.7e24 FLOPs on H800s
teortaxesTex · x · 2026-07-22
A compute back-of-the-envelope shows a frontier training run may have used about 3.7e24 FLOPs on H800s.
- The poster compares two ways of estimating the same run: one based on 0.385 PFLOPs × 2048 H800s × 55 days, the other using 37e9 × 14.8e12 × 6.
- The two numbers come out close, at 3.74e24 vs 3.28e24.
- The mismatch is attributed to the share of bf16 compute.
- The joke ending: “50K Hopper bros in shambles,” implying the scale is eye-watering even by current accelerator standards.
More from Infra
- Azure Architecture Diagram Builder adds MCP support for agent-driven Bicep workflows — davemccollough · 2026-07-22
- Engy posts live inference prices as Qwen3.6 undercuts GLM-5.2 on cached input — markjeffrey · 2026-07-22
- A hybrid local-plus-cloud inference model is the AI equivalent of 65 MPH driving — dmitry140 · 2026-07-22
- Alphabet capex call may matter less than what the spending is buying — tengyanAI · 2026-07-22
- U.S. firms may need Chinese models for cyber defense, yet one breach could trigger a ban — natolambert · 2026-07-22
- DeepSeek’s hardware quadrant would be notable, and Alibaba also has its own chips — teortaxesTex · 2026-07-22