DeepSeek-V4-Flash Local Inference Optimized: 27% Speedup on a Single DGX Spark
antirez · x · 2026-08-09
Developers have successfully optimized the local inference performance of DeepSeek-V4-Flash on a single DGX Spark using the DwarfStar engine, achieving a 27% speedup with the same weights.
- Decode speed: Increased from 22.5 to 28.6 tok/s
- Prefill speed: Steady at 1060 tok/s
- Optimization details: DwarfStar v0.5.6 added Spark-specific fast paths and enabled speculative decoding during the thinking phase, reaching a 74% draft acceptance rate.
This allows a 284B parameter model to run entirely on a desk without cloud dependency.
More from Infra
- AI Data Centers End Decades of Stagnant US Power Use, Challenging Anti-Growth Mindsets — AndyMasley · 2026-08-09
- Amazon's New Texas Data Center Power Plant Could Become a Top US Polluter — The Verge AI · 2026-08-09
- Upgrading to PyTorch 2.13 and CUDA 13 Doubles Video Resolution on Blackwell GPUs — Chemical-Bicycle3240 · 2026-08-09
- Training 10T Parameter Models Requires Over 50k GB300 GPUs — AccBalanced · 2026-08-09
- Amazon Data Centers Become the Biggest Pollution Source in the US — geox · 2026-08-09
- Matt Turck: Tech Industry Shouldn't Ignore Resistance to AI Data Centers — mattturck · 2026-08-09