Running 510GB DeepSeek V4.1 Flash on one DGX Spark: 113.6GB base plus 40MB domain sidecars
Physical_Toe_2499 · reddit · 2026-10-06
A developer demonstrated YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash (originally 510GB) on a single NVIDIA DGX Spark (GB10, 128GB unified memory) using a three-file architecture:
- Quantized base: VQ-8 scheme where every 8 consecutive expert weights become one 12/13-bit codebook index; codebooks trained per layer and shared across all 384 experts, stored in FP8. Zero-corpus, 113.6GB, a complete model on its own.
- Domain sidecars: 40MB each for finance, code, law, medicine, and science. They store per-channel multiplicative gains (FP4) plus router biases, solved layer-by-layer to reproduce the original MoE block outputs on domain text.
- Experimental post-training file: turns preferences into linear equations on last-layer expert gains; flips decision points from 55% to 88% on training requests but doesn't generalize yet.
Results: domain sidecars improve top-1 agreement vs. the original by 2.8–3.7 points (e.g., code 78.61%→82.26%), and English WikiText-2 actually improves slightly with any sidecar attached. Speculative decoding hits 43 tok/s on a real 14.1k-token agent request; prefill reaches 1,055 tok/s on 12.5k prompts. Decode runs at 82% of the GB10's 235GB/s memory-bandwidth wall. Limitations: teacher-forced metrics only, no human evaluation.
More from Infra
- Disco a dangerous short after HBM despec news; H2 China demand may flip — zephyr_z9 · 2026-10-06
- TensorFold bonds dual Thunderbolt 5 links for 83% throughput boost on Apple Silicon — AIFlow_ML · 2026-10-06
- Dual RX 6800 running Qwen3 27B: Linux + ROCm beats Windows Vulkan, 45 tok/s TG — SaGa31500 · 2026-10-06
- DeepSeek open-sources DeepGEMM, a clean and efficient GPU BLAS kernel library, now 8.5k stars — deepseek-ai · 2026-10-06
- AMD's Lisa Su predicts "very high" chip demand for years despite AI safety concerns — AryHHAry · 2026-10-06
- Photonic AI chip can help calculate its own corrections via in-situ gradient descent — bravo_abad · 2026-10-06