DeepSeek V4-Flash Slashes Compute to 66K GPU-hours Per 1T Tokens
teortaxesTex · x · 2026-08-01
A thread details the evolution of compute consumption across DeepSeek's models: dropping from 300K H800 GPU-hours per 1T tokens for V1, to 173K for V2, 180K for V3, and an estimated 66K for the latest V4-Flash.
For inference, V4-Flash and V2 have the same input/output prices, but cache hits are 5x cheaper and speed has roughly tripled or quadrupled. Their current best model is now smaller and cheaper than the version from 26 months ago.
Related event: DeepSeek Drastically Reduces Training Compute Costs Across Models(2 posts)→
More from Infra
- Cut Your $200/Month AI Bill: 10 Steps to Local AI Deployment — jason_mayes · 2026-08-01
- Local MXFP4 Testing Reveals Inconsistent Quality Among OpenRouter DS4 Flash Providers — antirez · 2026-08-01
- DeepSeek V4 Flash local benchmark nearly matches top frontier models from 5 months ago — joorklee · 2026-08-01
- New Method Pre-routes MoE Layers to Optimize I/O for Edge Streaming — dai_app · 2026-08-01
- 5TB of Data Stored on a Tiny Glass Slab Marks Microscopic Storage Breakthrough — TansuYegen · 2026-08-01
- antirez Enables Lossless MXFP4 Local Inference for DeepSeek v4 Flash — antirez · 2026-08-01