DeepSeek-V4-Flash Leaks: $2 Input / $6 Output, Runs on Single RTX 3090
Boring_Aioli7916 · reddit · 2026-08-03
A Reddit user revealed the pricing for the rumored DeepSeek-V4-Flash: $2 per 1M input tokens and $6 per 1M output tokens. Additionally, users have successfully run the IQ2XS quantized version locally on a single RTX 3090.
Related event: DeepSeek V4-Flash: 105x Cheaper but Stability Concerns(17 posts)→
More from Models
- Big Tech's GPU Hoarding Raises Open-Weight Hosting Costs Near Closed-Model Levels — xiaosun86 · 2026-08-03
- “Leading Models Still Make Stupid Mistakes”: Dev Spots Obvious Bug Missed by AI Final Audit — CtrlAltDwayne · 2026-08-03
- DeepSeek's Cache Read Pricing is 10x Cheaper Than Competitors — alejandroll10 · 2026-08-03
- Abacus AI Launches RouteLLM API for Task-Based and Cost-Aware Model Routing — bindureddy · 2026-08-03
- Open Weights Dominate: 9 of Top 10 Models on OpenRouter — zainhas · 2026-08-03
- OpenAI's Internal Astra Model Reportedly Solves 10 Major Math Problems, Breaking LLM Limits — ValerioCapraro · 2026-08-03