Single DGX Spark Decodes 15K Tokens from 124B Model with Stable Throughput
AcanthisittaOk1699 · reddit · 2026-08-14
A developer tested the community Q5 GGUF quantization of Ling-3.0-flash (124B parameters) on a single DGX Spark. Given a 33-token prompt, the model generated a massive 15,128 tokens in a single response, taking about seven minutes.
The most notable observation is the decoding throughput stability. Even as KV cache accumulated over 15,000 tokens, the decode speed remained remarkably flat, barely shifting from 35.62 tok/s to 35.68 tok/s throughout the entire generation process.
Related event: Single DGX Spark Tests 124B Model with High-Speed Decoding(2 posts)→
More from Infra
- NVIDIA Releases AI Tokenomics White Paper: A Framework for Monetizing Inference — nvidia · 2026-08-14
- The New Frontier of CS: Insane Potential of Shared Memory Across Computers — lauriewired · 2026-08-14
- OpenAI Previews Ultrafast Mode: GPT-5.6 Sol Hits 14x Speeds — OpenAI · 2026-08-14
- Detecting Performance Regressions Using ML and Hardware Counters — ZeroDark_Hereford · 2026-08-14
- Polymarket Prices Nvidia at 73% Chance to Be World's Largest Company by 2026 — Polymarket · 2026-08-14
- Debunking Seven Common Myths About AI Data Centers' Environmental and Economic Impact — Dan_Jeffries1 · 2026-08-14