Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD

HankYeomans · x · 2026-10-08

Developer HankYeomans shares a hybrid setup for running DeepSeek v4.1 Flash MXFP4 locally: the model weighs in at 456GB total, and he fits it using two GPUs (192GB VRAM combined) for the core, offloading experts that don't fit to system memory and SSD. He says response quality is now "pretty decent," while noting it's still a small test. The post illustrates a tiered-storage approach that lets enthusiasts run very large MoE models on limited VRAM.

Original post →

More from Infra

Infra channel →