RDNA4 local inference hits ~100 tok/s running Qwen3.8 Flash on dual R9700
Public_Umpire_1099 · reddit · 2026-09-14
A developer shipped an R9V update pushing local RDNA4 inference to 100 tok/s text generation with Qwen3.8 Flash Next IQ4XS on dual R9700 GPUs plus 128GB RAM, and added Q4KXL support at 50 tok/s.
Key fixes and notes:
- Crashes resolved via n-gram SSD streaming, with improved diagnostics and pinned image support
- Torture-tested for 12 hours with no instability found; Q4KXL still needs more tuning
- The author is also building an inference engine and a 6-month-old deep research/site-builder app, and will return to R9V after shipping those
- Community members in the Launch80 Discord report even better numbers with other quants, suggesting these non-fully-VRAM-resident configs are nearing their ceiling
More from Infra
- Perplexity launches Hybrid Compute to split AI tasks between cloud and local Mac — Aiden_Tech_Ai · 2026-09-14
- 4 sink tokens + 64-token window matches distilled linear attention, no training needed — burny_tech · 2026-09-14
- Anthropic reportedly signed a $13.7B, 6-year compute deal with RUM Group's Georgia site — rohanpaul_ai · 2026-09-14
- AI-written OpenSCAD + dual 12-inch fans fix NVIDIA Thor Dev Kit thermal throttling — catplusplusok · 2026-09-14
- Memory wall chart: compute up ~3x per two years, HBM bandwidth under 2x — Summit-Star001 · 2026-09-14
- Data centers are crowding out US private construction, WSJ data shows — GregCook2011 · 2026-09-14