DeepSeek V4-Flash Slashes Compute to 66K GPU-hours Per 1T Tokens

teortaxesTex · x · 2026-08-01

A thread details the evolution of compute consumption across DeepSeek's models: dropping from 300K H800 GPU-hours per 1T tokens for V1, to 173K for V2, 180K for V3, and an estimated 66K for the latest V4-Flash.

For inference, V4-Flash and V2 have the same input/output prices, but cache hits are 5x cheaper and speed has roughly tripled or quadrupled. Their current best model is now smaller and cheaper than the version from 26 months ago.

Related event: DeepSeek Drastically Reduces Training Compute Costs Across Models(2 posts)→

Original post →

More from Infra

Infra channel →