DeepSeek-V4-Flash Runs at 400 tps for Just $10/hour on Inference Endpoints

ben_burtenshaw · x · 2026-07-31

A developer shared performance stats for the DeepSeek-V4-Flash-0731 model on inference endpoints: it achieves a generation speed of 400 tokens/s while costing only $10 per hour to run. This demonstrates the extreme cost-efficiency and inference optimization of current open-source models.

Original post →

More from Infra

Infra channel →