Serving Qwen3.8-Flash-Next at 262K context on 8 RTX 3090s hits 1,200 tok/s

QuixiAI · x · 2026-09-06

Eric Hartford (QuixiAI, creator of Dolphin and Samantha) details how he served Qwen3.8-Flash-Next — a 125B-parameter hybrid of Gated DeltaNet linear attention, Qwen Sparse Attention, n-gram embedding tables and a built-in speculative drafter — on eight five-year-old RTX 3090s with no FP8 or NVLink, at its native 262,144-token context with exact bf16 KV cache.

Measured benchmarks (1,000 in / 2,000 out per request):

A 512 GiB pinned-RAM second tier lets evicted conversations resume in under a second instead of re-prefilling. His first working config managed only 409.7 tok/s and used quantized KV that was silently corrupting answers; the final setup delivered a net 2.9x speedup while doubling context and removing the harmful quantization. Full configs are open-sourced.

Original post →

More from Infra

Infra channel →