DeepSeek reportedly running four V4 model versions, user says culling is coming to cut inference costs
teortaxesTex · x · 2026-09-09
User teortaxesTex argues DeepSeek consistently optimizes for minimal inference cost, which rewards running as few model types as possible. He lists four currently live versions — V4-Pro-0813, V4-Flash-0731, V4-Flash-Vision-Exp, and V4.1 — and predicts some will be culled. Pushing back on complaints, he notes DeepSeek isn't AWS: users have no long-term enterprise contracts and no guarantee their use case will be preserved. Unverified speculation: no official source for the version names or any deprecation plan.
More from Infra
- IQE CEO: Indium Phosphide Substrate Supply a Key Risk as China Export Controls Bite — pstAsiatech · 2026-09-09
- Is a GTX 1660 Super 6GB enough to run ComfyUI locally? — srthk97 · 2026-09-09
- Running 285B DeepSeek-V4-Flash-Vision on 12x RTX 3090s at 120+ tok/s — ciprianveg · 2026-09-09
- After the OpenAI Debacle, a Case for Trusted Execution Environments for Frontier LLM Inference — Michael_D_Moor · 2026-09-09
- Running a $10,000 AI model at home: Fireworks AI engineer's agentic workflow — David Ondrej · 2026-09-09
- NVIDIA reportedly buying Hugging Face for $12.9B; ZipNN research cuts model storage ~20% — LChoshen · 2026-09-09