Single DGX Spark Tests 124B Model with High-Speed Decoding

Benchmark tests on a single DGX Spark reveal that the 124B Ling-3.0-flash model achieves 38.7 tok/s with almost no decoding speed loss during quantization, thanks to its MoE architecture.

2026-08-12 ~ 2026-08-13 · 2 related posts