Ling-3.0-flash MXFP4 Runs Locally on DGX Spark: 80 tok/s Decoding
niacolhealth · reddit · 2026-08-05
The MXFP4 quantized version of Ling-3.0-flash has been successfully deployed locally on a single DGX Spark.
Benchmark results show strong performance:
- Decoding speed: 80 tok/s
- Long-input prefilling: 2,500–3,500 tok/s
- Concurrency: Smoothly supports 3-4 concurrent users
This setup is suitable for private on-device inference, coding assistance, agent deployment, and offline batch processing.
More from Infra
- Running a Complex AI Task Costs Only $0.0042 — victormustar · 2026-08-06
- Opinion: AI and Robotics Will Make Labor Abundant, Compute and Energy Are the Next Oil — VraserX · 2026-08-06
- Dassault Systèmes Partners with NVIDIA to Accelerate Simulation via AI — NVIDIAAI · 2026-08-06
- Ezra Klein Podcast Discusses AI Compute Crisis and Data Center Moratoriums — kevinsxu · 2026-08-06
- Running MiniMax H3 on RTX 5090: Node Optimization Triples Speed — WARRIORPSIX · 2026-08-06
- Opinion: AI Data Centers Are the Future, Canada Must Overcome Backlash — LoganGrasby · 2026-08-05