DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s

AccBalanced · x · 2026-09-12

A developer demonstrated DeepSeek-V4.1-Flash generating tokens on one RTX 5090 with 125.7 GiB RAM, while most of a 502GB GGUF lives on NVMe. Hot experts stream through VRAM/RAM, cold experts stay on SSD, and a 196B Engram memory is disk-backed. Key numbers: 552B backbone, only 8B active per input token and 16B per output token, 5.12 tok/s on new content (up to 21.27 tok/s with resident data), and 0.9967 logit correlation to the reference model. Commentators see this — like Qwen3.8-Flash-Next — as evidence model capacity is decoupling from active compute, a blueprint for running far larger models locally.

Original post →

More from Infra

Infra channel →