Serving Qwen3.8-Flash-Next at 262K context on 8 RTX 3090s hits 1,200 tok/s
QuixiAI · x · 2026-09-06
Eric Hartford (QuixiAI, creator of Dolphin and Samantha) details how he served Qwen3.8-Flash-Next — a 125B-parameter hybrid of Gated DeltaNet linear attention, Qwen Sparse Attention, n-gram embedding tables and a built-in speculative drafter — on eight five-year-old RTX 3090s with no FP8 or NVLink, at its native 262,144-token context with exact bf16 KV cache.
Measured benchmarks (1,000 in / 2,000 out per request):
- 1 concurrent: 140 tok/s, 14s median latency
- 8 concurrent: 590-600 aggregate tok/s
- 32 concurrent: peak 1,199.8 tok/s
A 512 GiB pinned-RAM second tier lets evicted conversations resume in under a second instead of re-prefilling. His first working config managed only 409.7 tok/s and used quantized KV that was silently corrupting answers; the final setup delivered a net 2.9x speedup while doubling context and removing the harmful quantization. Full configs are open-sourced.
More from Infra
- NanoDiffuser: sub-400 MiB text-to-image model running in-browser via WebGPU — jason_mayes · 2026-09-06
- Wall Street veteran: the world's biggest data center consumes water like just 3 golf courses — rohanpaul_ai · 2026-09-06
- Japan says $550B U.S. investment pact advancing, AI and semiconductors to play major role — Polymarket · 2026-09-06
- QuixiAI open-sources SlimServe, the inference stack behind the 8x3090 Qwen deployment — QuixiAI · 2026-09-06
- T-Glass shortage worsens: Kinsus losing 10-15% of monthly ABF revenue, 25% capacity expansion planned for 2027 — zephyr_z9 · 2026-09-06
- Bump-less 3D stacking goes practical: Intel Diamond Rapids first, AMD Zen rumored next — bookwormengr · 2026-09-06