Benchmark: 124B Model Hits 38.7 tok/s on a Single DGX Spark
AcanthisittaOk1699 · reddit · 2026-08-13
An independent tester, sudoingX, spent a week benchmarking a 124B parameter model on a single DGX Spark.
Key benchmark results:
- Official INT4: Achieved 38.7 tok/s when configured correctly, marking the fastest path found.
- Community GGUF: Reached 35.2 tok/s.
- Comparison: This is 2.4x faster than what DeepSeek V4 Flash does on the same machine.
The tester initially claimed the official quants wouldn't run on a single Spark, but later corrected himself after finding the right configuration, stating that Ling-3.0-flash earned a permanent seat on his box.
Related event: Single DGX Spark Tests 124B Model with High-Speed Decoding(2 posts)→
More from Infra
- Mistral Pivots to European Inference Provider, Leveraging Sovereign Infra Over Weaker Models — teortaxesTex · 2026-08-13
- AI Boom Spreads: Investors Target Chip Fab and Data Center Suppliers — Polymarket · 2026-08-13
- Extreme Optimization: Running 33B Video Generation Model on M4 ANE — antirez · 2026-08-13
- LiteLLM Supply Chain Attack Leaks 153GB from 2,488 Orgs Including Nvidia and AWS — wunderwuzzi23 · 2026-08-13
- Wetty: Run a Terminal Emulator in Your Browser via Node.js and SSH — tom_doerr · 2026-08-13
- SK Hynix to Invest $38.1B in Two New Memory Fabs in Korea — Beth_Kindig · 2026-08-13