R9V: Custom RDNA4 Kernels Boost Qwen3.8 Throughput by 30x
Public_Umpire_1099 · reddit · 2026-08-31
The author released R9V, a set of highly custom inference kernels tuned specifically for AMD's RDNA4 (R9700s) architecture, integrated into vLLM-Radiance to significantly boost performance for Qwen3.8-Flash-Next and Muse Glimmer 30B.
Key Benchmarks (Dual R9700, 128GB RAM):
- Qwen3.8 Flash Next: PP8192 hit 1510 tok/s (30x increase), TG256 hit 78 tok/s (3x increase).
- Muse Glimmer 30B: Outperformed llama.cpp (ROCm/Vulkan) across the board.
Technical Details:
- Deep utilization of RDNA4's Wave32 design, DPP operations, and integer dot instructions.
- Optimized MTP (Speculative Decoding) by reusing token routes for hot experts, achieving 27% bandwidth optimization.
- Optimized Prefill by grouping prompt tokens by expert (group size 16).
- For dense models, implemented effective reuse of multivector weights and HyperConnection fusion.
More from Infra
- Exclusive: SK hynix weighs Intel Foundry for next-gen HBM4E base dies, breaking TSMC dependence — BenBajarin · 2026-08-31
- Parody: The $200/mo user costing OpenAI $14,000/mo — BuildersReadOnAI · 2026-08-31
- Measuring what Windows apps expose to computer-use agents — Frequent-Ad-836 · 2026-08-31
- Rayrun Implements sPTC to Speed Up AI Responses by 20% — lucgagan · 2026-08-31
- Dev asks if HY4's 1.25-bit quantization (1.5TB→200GB, 98% retention) is worth porting to Qwen3.8-Flash-Next — TemperatureOk3561 · 2026-08-31
- Samsung Takes Lead in HBM4 as SK Hynix and Micron Struggle — AccBalanced · 2026-08-31