Running DeepSeek-V4-Flash (284B MoE) at 75 tok/s on 2× DGX Spark: Full Recipe & 11 Gotchas
Striking-Swim6702 · reddit · 2026-08-11
A developer shares a comprehensive guide on deploying DeepSeek-V4-Flash-0731 (284B MoE) in production across two interconnected DGX Sparks.
- Performance: Achieves 74.8 tok/s on real coding prompts with a 42s TTFT on 64k prompts, backed by a 2.74M token KV capacity for full 1M context.
- Gotchas: Details 11 hidden traps including dual-rail RDMA caps, Ray OOM-killing workers due to unified memory quirks, missing env vars causing garbled Chinese output, and speculative decoding inversion under concurrency.
- Integration: Adapted Codex CLI to the new Responses API and successfully ran agentic evaluations implementing a multi-file concurrency feature in a 1,831-file codebase.
Related event: DeepSeek V4 Flash Excels in Dual DGX Spark Cluster Tests(3 posts)→
More from coding & agent
- Ask Your AI Agent for a Markdown Checklist, Not Just a Plan — JnBrymn · 2026-08-12
- AI Code Quality Depends on the Constraints You Set Around Agents — rseroter · 2026-08-12
- Payments and Service Discovery for Autonomous Agents: No Standard Yet — Nata_Elisym · 2026-08-12
- Agent Success Rate Drops on Repeat: Paper Reveals Computer Use Reliability Trap — xwang_lk · 2026-08-12
- Cursor reportedly set to launch Composer 3 as users await evals — kimmonismus · 2026-08-12
- Developer Showcases x40-Powered Agent Endpoint for Reverse Phone Lookup — MurrLincoln · 2026-08-12