Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks
NVIDIAAI · x · 2026-09-01
Benchmarked Qwen3.8 Flash Next on 2 DGX Sparks using NVFP4 format with 64 concurrent users. It generated 32,768 output tokens in 78.82 seconds, achieving a throughput of 415.7 tok/s. The test passed a 64K/user usable-context stress test with zero leakage, noted as a limit test rather than a production recipe.
More from Infra
- Google Cloud Monitoring MCP Connector Released — modelcontextprotocol · 2026-09-01
- Distributed.systems发布可审计的Agent基础设施 — arthurcolle · 2026-09-01
- Does enabling ChatGPT Memory or history reference increase token usage? — ssunki · 2026-09-01
- Engineer fixes ROCm inference crash on MI350X, uncovers 9 bugs in deep dive — AnushElangovan · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01