DeepSeek-V4-Flash Runs at 400 tps for Just $10/hour on Inference Endpoints
ben_burtenshaw · x · 2026-07-31
A developer shared performance stats for the DeepSeek-V4-Flash-0731 model on inference endpoints: it achieves a generation speed of 400 tokens/s while costing only $10 per hour to run. This demonstrates the extreme cost-efficiency and inference optimization of current open-source models.
More from Infra
- TSMC CoWoS Capacity Forecast to Exceed 2.3M Wafers by 2027, Significantly Revised Upwards — Beth_Kindig · 2026-08-01
- Tim Cook's Final Earnings Call: Memory Chip Shortage to Impact iPhone and Mac Supply — fortune · 2026-07-31
- Cerebras exec departs: wafer-scale chip breaks Moore's Law, future potential immense — draecomino · 2026-07-31
- More GB300s Online on LightningAI; Pangram 4 Hits 99% AI Text Detection Accuracy — LightningAI · 2026-07-31
- DeepSeek Rumored to Go All In on TileLang for Kernel Development — zephyr_z9 · 2026-07-31
- Gigawattonomics Model: AI Data Centers Pay Off Faster Than Expected — BenBajarin · 2026-07-31