From one 3090 to 20 DGX Sparks: a home local-LLM cluster epic, 2.8T Kimi K3 at 20 t/s
ciprianveg · reddit · 2026-10-04
A Redditor chronicles going from a single 3090 running LLaMA 33B to a 16-node DGX Spark cluster (shared with his brother) running 2.8T Kimi K3 — all on locally hosted models, never paying for a commercial API. Milestones: Threadripper + 512GB RAM to run DeepSeek 671B at 8 t/s; 16×3090 on a 100Gbit network for Qwen 397B (until house fuses and 6kW draw said no); then 4 linked ASUS GB10 units running Qwen 397B at 30 t/s on just 400W. He published the first working 8x and 16x Spark cluster solutions on NVIDIA forums, tuning Kimi K3 from an unusable 7 t/s at 100k context to 20 t/s at 300k via multiple vLLM/SGLang iterations. Next: 4 more Sparks so a GLM 5.3 Flash runs 24/7 alongside the big cluster (GLM 5.3 + MiMo 2.6 Pro, 16x Kimi K3, or Qwen 3.8 2.4T) — and 4 more for his younger brother.
More from Infra
- Running a 100B+ Qwen3 model locally on 64GB RAM: good vibes, short 192K context — lxfater · 2026-10-04
- TPU cost per million tokens beats NVIDIA Blackwell, giving Google an edge — cgarciae88 · 2026-10-04
- Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU — MoonsvnLyn · 2026-10-04
- Dev shares concurrency sweep method: TTFT, ITL and tok/s on 4xB200 for local models — TheZachMueller · 2026-10-04
- NVIDIA AIPerf docs go live: a package for performance-testing AI models — TheZachMueller · 2026-10-04
- One prompt freed 39.7 GB: Claude Code + ccmd MCP safely cleans dev caches — julsimon · 2026-10-04