Running DeepSeek-V4-Flash-Vision on dual RTX 6000: 350K context at 7 concurrent
DeedleDumbDee · reddit · 2026-09-03
The author upgraded their work LLM server from DSV4-Flash-0731 to the new Vision-EXP model, serving it via vLLM behind OpenWebUI.
Deployment details: after fixing KV cache memory issues, they reached 2.5M-token KV cache, 350K context at 7 concurrent requests, 120 tok/s generation and 3,000–7,000 tok/s prefill.
The test: a single self-contained HTML file using Three.js rendering an animated solar system—real orbital periods (0.24 to 164.8 Earth years), Kepler's equation solved numerically, real eccentricities/inclinations/axial tilts, NASA-confirmed moon counts (Jupiter 115, Saturn 293) rendered via InstancedMesh, plus orbit controls and speed UI. A useful reference for local deployment of a large vision model.
More from Models
- Codex, Claude and Grok All Went Down at the Same Time — fekdaoui · 2026-09-03
- Gemini has improved exponentially, Grok slower and often unavailable: Damodaran — pdamodaran · 2026-09-03
- Hacker upgrades stolen account to $200 ChatGPT Pro with no fraud checks from OpenAI or bank — CucharitaDePalo · 2026-09-03
- Muse Spark 1.3 Update Gains Attention, Users Urged Not to Miss — D3VAUX · 2026-09-03
- 64GB Mac users debate best sub-40B MoE: is Qwen-3.8-35B worth the wait? — chibop1 · 2026-09-03
- Microsoft's MAI-Transcribe-2 hits 2.0% WER at 411x real-time, priced at $1.67 per 1,000 minutes — ArtificialAnlys · 2026-09-03