DeepSeek-V4.1-Flash on 2x DGX Spark: TP2 Patches Cut First Token From 33.8s to 4.9s

rez0__ · x · 2026-10-12

An open-source repo publishes engine patches and benchmarks for serving DeepSeek-V4.1-Flash (EXL3 2.9bpw) with tensor parallelism across two NVIDIA GB10 nodes (DGX Spark). Results: single-stream decode 90→99 tok/s (code) and 56→61 tok/s (prose), 4 streams at 107→115 tok/s total, prefill +5-6%, and agent replay first token from 33.8s down to 4.9s.

Based on 24h of real coding-agent traffic (1,853 requests, median prompt 163K tokens, 80% cacheable), the prompt-cache analysis shows: 24% of prefilled tokens resumed from kept state, 40% never hit cache due to image-carrying requests, 34% failed to resume because kept state was overwritten in the shared KV pool, and only 2% were real misses — roughly three quarters of all prefill time redid work that had already been done.

Original post →

More from coding & agent

coding & agent channel →