Running 510GB DeepSeek V4.1 Flash on one DGX Spark: 113.6GB base plus 40MB domain sidecars

Physical_Toe_2499 · reddit · 2026-10-06

A developer demonstrated YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash (originally 510GB) on a single NVIDIA DGX Spark (GB10, 128GB unified memory) using a three-file architecture:

Results: domain sidecars improve top-1 agreement vs. the original by 2.8–3.7 points (e.g., code 78.61%→82.26%), and English WikiText-2 actually improves slightly with any sidecar attached. Speculative decoding hits 43 tok/s on a real 14.1k-token agent request; prefill reaches 1,055 tok/s on 12.5k prompts. Decode runs at 82% of the GB10's 235GB/s memory-bandwidth wall. Limitations: teacher-forced metrics only, no human evaluation.

Original post →

More from Infra

Infra channel →