Running LLMs on DGX Spark: DeepSeek V4 Flash Tops the Ranking
andersonbcdefg · x · 2026-08-03
A developer shared real-world benchmarks and rankings of serious LLMs running locally on an NVIDIA DGX Spark (128GB RAM).
- #1: DeepSeek V4 Flash 0731. Using 3-bit quantization, it runs at 16.5 tok/s. Although it's the slowest, it's the biggest brain (284b total / 13b active params) that fits into 128GB, making it a pro thinker worth the wait for hard problems.
- #2: Laguna S 2.1. Using NVFP4, it runs at 19.8 tok/s (up to 37 tok/s on code). It is the most trusted model for real agent work, capable of opening a todo list, tracking steps, and maintaining context over long hours without losing the thread.
However, the original thread was quote-tweeted by another user who joked that this complicated setup process is actually a fantastic advertisement to never buy a DGX Spark.
Related event: DeepSeek V4 Flash Local Deployment Benchmarks Revealed(3 posts)→
More from Infra
- Nemotron 3 Nano Omni Hits 264 tok/s Native on DGX Spark — ivan_bezdomny · 2026-08-03
- Global AI compute to hit 200M H100-equivalents by 2028, fueling agentic loop toward ASI — 新智元 · 2026-08-03
- tinybox Dual-GPU Edition Hits 245 tok/s Running DeepSeek — AccBalanced · 2026-08-03
- Troubleshooting KV Cache Misses Caused by Multiple Agent Tool Calls — CentrifugalMalaise · 2026-08-03
- Wafer serves Kimi K3 on AMD MI355X with 3.8x throughput and 71% lower cost vs B200 — SumitGup · 2026-08-03
- Amazon Graviton Revenue Commitments Surge Nearly 3x QoQ — Beth_Kindig · 2026-08-03