RDNA4 local inference hits ~100 tok/s running Qwen3.8 Flash on dual R9700

Public_Umpire_1099 · reddit · 2026-09-14

A developer shipped an R9V update pushing local RDNA4 inference to 100 tok/s text generation with Qwen3.8 Flash Next IQ4XS on dual R9700 GPUs plus 128GB RAM, and added Q4KXL support at 50 tok/s.

Key fixes and notes:

Original post →

More from Infra

Infra channel →