462GB DeepSeek model runs on two desk-side DGX Sparks with experts squeezed to 2.77 bits
Teknium · x · 2026-10-06
Developer sudoingX loaded the rumored DeepSeek v4.1 flash onto two DGX Sparks using @0xSero's recipe (version unverified):
- The 462GB checkpoint spans two boxes; routed experts squeezed to 2.77 bits with EXL3 and engram tables stream straight off NVMe, so a 462GB model fits in 256GB of total memory, up in 15 minutes over a ConnectX cable
- 262k context with dspark k=7 drafting on a single stream: 34.8 tok/s on code, 22.5 tok/s on prose, slightly below the original author's 41/29
- Serving was retuned for agent work: leaner KV cache to leave room for build tools, browser and tests, plus locked earlier builds so the agent can't cheat
- Paired with a Hermes agent harness; text, tools and vision all passed smoke tests
It doubles as a reproducible dual-box desktop inference recipe: source-built vLLM fork with exllamav3 and b12x kernels, NCCL pointed at the RoCE port, and more.
More from coding & agent
- Dev builds native Mac app to message multiple AI agents simultaneously — msg · 2026-10-06
- Dev uses GitHub Copilot app to generate a rebuildable 60s hype video via PR — DanWahlin · 2026-10-06
- Oxylabs raises $130M at $3.6B valuation, first outside capital, betting on AI agents — Ronangmi · 2026-10-06
- PM open-sources GitDocket: manage AI coding epics, tasks and wikis in Git — bjorgen · 2026-10-06
- Amelia Wattenberger: AI is still in its terminal era, and GUI history holds lessons — Wattenberger · 2026-10-06
- Ex-Shopify principal designer Eric Hu joins Cognition as principal designer — Kyrannio · 2026-10-06