Qwen3.8-Flash-Next on a single R9700: 863 t/s prefill and 35 t/s decode at 230k context
Designer_Elephant227 · reddit · 2026-10-02
A Redditor shared full details of running Qwen3.8-Flash-Next locally on a single AMD R9700 via the exllamav3 rocm fork, asking whether the setup is optimal.
- Model: turboderp's exl3 5.05 bpw quant (6-bit head/vision, 5-bit mtp)
- Expert offload: 112 of 512 experts per layer on GPU, 400 on CPU (16 threads); a 102GB bf16 ngram table streamed from NVMe
- 262k context, q8 KV cache, chunk size 4096, MTP drafting enabled (53-54% acceptance)
- Measured at 230k context prose: 863 t/s prefill (101k new tokens, rest from prefix cache), 34.7 t/s decode
- Required patching one file for gfx12
More from Infra
- The Neocloud Reality Check: Why Your Next AI Project May Skip the Big Clouds Entirely — DavidLinthicum · 2026-10-02
- A 3-step guide to open models: picking, hosting locally or via OpenRouter — every · 2026-10-02
- Community inference contest hits 1900 tps prefill, 2.5x faster than omlx baseline — HankYeomans · 2026-10-02
- San Antonio district hosts a dozen data centers as industry camouflages them in woods — The Verge AI · 2026-10-02
- Need a GPU fast? Self-serve clouds like RunPod offer instant spin-ups without contracts — DavidLinthicum · 2026-10-02
- Amazon writes 3,000-word blog warning communities not to block data centers — The Verge AI · 2026-10-02