DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s
AccBalanced · x · 2026-09-12
A developer demonstrated DeepSeek-V4.1-Flash generating tokens on one RTX 5090 with 125.7 GiB RAM, while most of a 502GB GGUF lives on NVMe. Hot experts stream through VRAM/RAM, cold experts stay on SSD, and a 196B Engram memory is disk-backed. Key numbers: 552B backbone, only 8B active per input token and 16B per output token, 5.12 tok/s on new content (up to 21.27 tok/s with resident data), and 0.9967 logit correlation to the reference model. Commentators see this — like Qwen3.8-Flash-Next — as evidence model capacity is decoupling from active compute, a blueprint for running far larger models locally.
More from Infra
- TensorSharp hits 41 tok/s decoding DeepSeek V4.1 Flash on 8× A40 — fuzhongkai · 2026-09-12
- What actually runs AI models at the edge in 2026: Mac mini, DGX Spark, iPhone 17 Pro — MaziyarPanahi · 2026-09-12
- Running Qwen3.8 Flash Next on dual RTX 3090: full llama.cpp config shared for tuning — ChopSticksPlease · 2026-09-12
- UAE redesigns 5GW AI campus with bunkers and air defenses after Iranian strikes on Gulf cloud facilities — mark_k · 2026-09-12
- Running 100-200 agents daily: disk space is now the bottleneck, not compute — vincent_koc · 2026-09-12
- Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96% — AccBalanced · 2026-09-12