Running a 510GB DeepSeek model on a 128GB DGX Spark by pruning unused MoE experts
pbaylies · x · 2026-09-14
@0xBakeer demonstrates running the 510GB DeepSeek-V4.1-Flash with 256k context and thinking mode on a single 128GB DGX Spark — no smaller model, no cloud.
The trick exploits MoE sparsity: each token activates only 6 of 384 experts, so he pre-selects which experts the target job needs and loads only those.
More from Infra
- Sandbox tip: bake dependencies into the image instead of pip-installing at runtime — xeophon · 2026-09-14
- US's No.2 law firm Latham & Watkins builds in-house AI stack with Nvidia servers — ayushtweetshere · 2026-09-14
- Oracle Cuts Double-Digit % of Some Teams While Hiring Aggressively for Data Centers and AI — mkheck · 2026-09-14
- Running Two Models Across Strix Halo + r9700 Hits OOM: Full Config Shared — El_90 · 2026-09-14
- Musk: AI will be 99% of SpaceX's value within four to five years — XFreeze · 2026-09-14
- Weaker enterprise HBM demand could finally normalize DRAM and NAND pricing, argues analyst — eyishazyer · 2026-09-14