Running DeepSeek V4 Flash 155G on DGX Spark: 2-bit Quantization & MTP Benchmarks

Puzzleheaded_Base302 · reddit · 2026-08-03

The author details the process and benchmarks of running the 155 GB DeepSeek-V4-Flash-0731 locally on a single DGX Spark (GB10, 121.7 GiB unified memory) using vLLM-Moet with 2-bit quantization.

Key Performance Metrics:

Critical Deployment Gotchas:

Original post →

More from Infra

Infra channel →