Running a 124B model on one 128GB desktop GPU: the engineering story behind the benchmark
nikola_mr64990 · x · 2026-09-18
A community author combed through a 70-post NVIDIA developer forum thread on running ling-3.0 flash (124B params) on a single DGX Spark with 128GB memory — and the real story beats the int4/fp4 benchmark screenshots.
- The heavy lifting came after the first successful run: quantization, custom kernels, MTP, KV cache tuning, memory profiling, and a series of vLLM fixes.
- The writeup combines official numbers with the author's own measurements, showing how open-source collaboration turned a headline claim into a working local deployment.
- Takeaway: open source moves differently — the model release is just the starting point; community engineering makes it usable.
Related event: Community Gets 124B Ling-3.0 Flash Running on a Single 128GB DGX Spark(3 posts)→
More from Infra
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- PlanetScale's TIN beats Postgres GIN full-text search: 212ms vs 288s at p99 — DanielLockyer · 2026-09-18
- NVIDIA shows 100x faster scikit-learn spectral clustering with cuML — NVIDIA Developer · 2026-09-18
- ChatGPT desktop app leaks context to the cloud by default — here's how to swap in Ollama — Technovangelist · 2026-09-18
- Reddit thread: what max-context KV reservations actually cost beyond concurrency — werunm · 2026-09-18
- Dev Patches vLLM to Run DiffusionGemma, Live Evals Show It Ties on Smarts but Loses on Speed to APIs — bodonoghue85 · 2026-09-18