DGX Spark Runs Qwen3.6-35B for 64 Concurrent Users
StefanoGogioso · x · 2026-07-10
Running Qwen3.6-35B on a single DGX Spark:
- Supports 64 concurrent users.
- Throughput reaches 700+ tok/s.
- Each user retains their own prompt and KV cache, using vLLM to batch all active streams step-by-step into the GPU.
- The post includes a link to the implementation recipe, noting the total system power draw is around 38W.
More from Infra
- Vercel AI Gateway data shows Anthropic, OpenAI and Google at 97.09% spend share — cramforce · 2026-07-21
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21
- A shopping app demo ties OpenTelemetry, Dynatrace and Port into agentic ops — Pavan_Belagatti · 2026-07-21