DeepSeek-V3 Trained With Only 180K GPU-Hours, Slashing MoE Compute Costs

teortaxesTex · x · 2026-08-01

A comparison of DeepSeek's generational training compute reveals massive efficiency gains: V1 required 300K H800 GPU-hours per 1T tokens, while V2 and V3 dropped to 173K and 180K respectively. The author estimates their next model (V4-Flash) might need only around 66K hours, totaling roughly 2M GPU-hours.

This implies DeepSeek's current best model is smaller and cheaper to train than its version from 26 months ago. The author praises this as the ultimate validation of Scaling laws for MoEs, contrasting DeepSeek's transparent inference unit economics with the secretive myth-making of Western frontier AI companies.

Related event: DeepSeek Drastically Reduces Training Compute Costs Across Models(2 posts)→

Original post →

More from Infra

Infra channel →