antirez demos DGX Spark streaming a half-trillion-parameter model from SSD, "kinda usable for QA"
antirez · x · 2026-09-14
antirez shared a demo of a DGX Spark executing a half-trillion-parameter model via SSD streaming — "kinda usable for QA," though not for coding unless you let it run overnight. He admits README prefill is still much slower than it should be compared to Metal, an implementation issue he plans to fix in a next commit. A rare first-hand look at memory-streamed on-device inference of very large models from a high-profile developer.
Related event: antirez Runs 500B-Param DeepSeek on a Single DGX Spark via SSD Streaming(3 posts)→
More from Infra
- Chart shows hyperscalers went on a CAPEX spree after DeepSeek — and Kimi broke the logic — StewartalsopIII · 2026-09-14
- Free My VRAM: open-source tool auto-unloads idle ComfyUI models, 29.8GB to 0.8GB — WazaqG · 2026-09-14
- Nvidia Partners With Palantir to Chase the $500 Billion Sovereign AI Market — Beth_Kindig · 2026-09-14
- Meme Post Roasts Every Local LLM Camp: Apple, AMD, Nvidia and Offload Users — Saren-WTAKO · 2026-09-14
- No.2 US law firm Latham & Watkins buys Nvidia servers to fine-tune open models in-house — ai · 2026-09-14
- Andrew Chen's homelab for local AI: 5090 eGPU, dual DGX Spark and a routing plugin — andrewchen · 2026-09-14