Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD
HankYeomans · x · 2026-10-08
Developer HankYeomans shares a hybrid setup for running DeepSeek v4.1 Flash MXFP4 locally: the model weighs in at 456GB total, and he fits it using two GPUs (192GB VRAM combined) for the core, offloading experts that don't fit to system memory and SSD. He says response quality is now "pretty decent," while noting it's still a small test. The post illustrates a tiered-storage approach that lets enthusiasts run very large MoE models on limited VRAM.
More from Infra
- Splash 1.3.0 cuts local agent first-token time from 19s to 1s via SSD offloading — songhan_mit · 2026-10-08
- StackOverflow 2026 Survey: Postgres Stays #1 at 58%, Supabase Climbs to 8% — dshukertjr · 2026-10-08
- GB300 compute capacity bought with opencode credits in all-token deal — const_reborn · 2026-10-08
- Building an ultra-high throughput AI-SQL engine: cutting LLM call costs with query plans — HamelHusain · 2026-10-08
- Together is now the #1 provider by token volume on OpenRouter — NVIDIAAI · 2026-10-08
- umbrelOS 2.0 makes local AI a two-install setup: Ollama + Open WebUI with auto GPU detection — JosephJacks_ · 2026-10-08