Running 285B DeepSeek-V4-Flash-Vision on 12x RTX 3090s at 120+ tok/s
ciprianveg · reddit · 2026-09-09
Developer ciprianveg published a fully reproducible setup for running DeepSeek-V4-Flash-Vision-Exp (285B MoE) on consumer Ampere GPUs:
- FP4 experts + FP8 attention, 157 GB weights, SM86-compatible vLLM build
- 60+ tok/s decode on 10x RTX 3090 (TP2xPP5, 240W cap) with DSpark speculative decoding (k=3); 120+ tok/s on 12 GPUs (TP4xPP3)
- Vision, tool calls and spec decoding all working; 1M context without offload / 4M with RAM offload, 3,500 tok/s long-context prefill
Ships with a pre-built Docker image ghcr.io/ciprianveg/3090-vllm:dsv4-flash-vision-sm86 and a GitHub repo with build guide, start scripts and runtime patches covering DSpark propose-gate (spec+vision fix), scheduler mm×spec row-crossing, grammar-bitmask validation, ViT OOM fix, FlashInfer workspace-lane keying and more.
More from Infra
- IQE CEO: Indium Phosphide Substrate Supply a Key Risk as China Export Controls Bite — pstAsiatech · 2026-09-09
- Is a GTX 1660 Super 6GB enough to run ComfyUI locally? — srthk97 · 2026-09-09
- DeepSeek reportedly running four V4 model versions, user says culling is coming to cut inference costs — teortaxesTex · 2026-09-09
- After the OpenAI Debacle, a Case for Trusted Execution Environments for Frontier LLM Inference — Michael_D_Moor · 2026-09-09
- Running a $10,000 AI model at home: Fireworks AI engineer's agentic workflow — David Ondrej · 2026-09-09
- NVIDIA reportedly buying Hugging Face for $12.9B; ZipNN research cuts model storage ~20% — LChoshen · 2026-09-09