From one 3090 to 20 DGX Sparks: a home local-LLM cluster epic, 2.8T Kimi K3 at 20 t/s

ciprianveg · reddit · 2026-10-04

A Redditor chronicles going from a single 3090 running LLaMA 33B to a 16-node DGX Spark cluster (shared with his brother) running 2.8T Kimi K3 — all on locally hosted models, never paying for a commercial API. Milestones: Threadripper + 512GB RAM to run DeepSeek 671B at 8 t/s; 16×3090 on a 100Gbit network for Qwen 397B (until house fuses and 6kW draw said no); then 4 linked ASUS GB10 units running Qwen 397B at 30 t/s on just 400W. He published the first working 8x and 16x Spark cluster solutions on NVIDIA forums, tuning Kimi K3 from an unusable 7 t/s at 100k context to 20 t/s at 300k via multiple vLLM/SGLang iterations. Next: 4 more Sparks so a GLM 5.3 Flash runs 24/7 alongside the big cluster (GLM 5.3 + MiMo 2.6 Pro, 16x Kimi K3, or Qwen 3.8 2.4T) — and 4 more for his younger brother.

Original post →

More from Infra

Infra channel →