DeepSeek v4.1 Flash costs under $3 for 500M tokens, hits 315 tok/s decode
gaganghotra_ · x · 2026-09-16
A user benchmarked DeepSeek v4.1 Flash via its API and found it remarkably cheap: over 500 million tokens cost less than $3.
On speed, they measured a peak decode rate of 315 tokens/s and prefill throughput of 52k tokens/s, making it both extremely affordable and fast.
More from Models
- Analyst argues Meta could actually exclude user data from training, unlike OpenAI's lawyerly wording — ivan_bezdomny · 2026-09-16
- DeepSeek V4.1 Flash Keeps Timing Out on 2-Hour Agentic Benchmarks, Author Shares Failure Logs — sebnadeau · 2026-09-16
- GPT-6 Astra remakes Game of Thrones in low-poly Blender after 8h autonomous run — OWazabi · 2026-09-16
- Why Meta Could Actually Keep Your Agent Data Out of Training — and Frontier Labs Can't — ivan_bezdomny · 2026-09-16
- TypeSafe AI launches Jev, a model for fast structured decisions with confidence scores — yogthinks · 2026-09-16
- GPT-6-Astra gets "depressed" in Minecraft after creeper blows up its base, farms potatoes for hours — scaling01 · 2026-09-16