DeepSeek-V4.1-Flash on 2x DGX Spark: TP2 Patches Cut First Token From 33.8s to 4.9s
rez0__ · x · 2026-10-12
An open-source repo publishes engine patches and benchmarks for serving DeepSeek-V4.1-Flash (EXL3 2.9bpw) with tensor parallelism across two NVIDIA GB10 nodes (DGX Spark). Results: single-stream decode 90→99 tok/s (code) and 56→61 tok/s (prose), 4 streams at 107→115 tok/s total, prefill +5-6%, and agent replay first token from 33.8s down to 4.9s.
Based on 24h of real coding-agent traffic (1,853 requests, median prompt 163K tokens, 80% cacheable), the prompt-cache analysis shows: 24% of prefilled tokens resumed from kept state, 40% never hit cache due to image-carrying requests, 34% failed to resume because kept state was overwritten in the shared KV pool, and only 2% were real misses — roughly three quarters of all prefill time redid work that had already been done.
More from coding & agent
- Custom church payment system: GPT security audit caught a critical flaw before it was too late — petergyang · 2026-10-12
- AI Agent Does CAD: KiCad MCP Reads Mounting Holes for a Robot Build — burhop · 2026-10-12
- Veteran dev: the future of apps is just telling your agent what you want — dreamwieber · 2026-10-12
- Cloudflare's Think proposes an execution ladder: LLMs pick their own sandbox per task — irvinebroque · 2026-10-12
- Opus 5.5 Compared: Claude Teammates Push Back, Codex Quietly Drifts Behind Green Tests — Sauers_ · 2026-10-12
- AI Writes Code Faster — So Why Aren't We Shipping Faster? — kristiyanstoyanovAI · 2026-10-12