Single DGX Spark Tests 124B Model with High-Speed Decoding
Benchmark tests on a single DGX Spark reveal that the 124B Ling-3.0-flash model achieves 38.7 tok/s with almost no decoding speed loss during quantization, thanks to its MoE architecture.
2026-08-12 ~ 2026-08-13 · 2 related posts
- Ling-3.0-flash Quantization Benchmarks: MoE Architecture Preserves Decode Speed — AcanthisittaOk1699 · 2026-08-12
- Benchmark: 124B Model Hits 38.7 tok/s on a Single DGX Spark — AcanthisittaOk1699 · 2026-08-13