Dual R9700 RDNA4 Setup Hits 5000 tok/s Prefill on Qwen 3.8 27B, Seeks Better Engines
N34257 · reddit · 2026-10-06
A user runs Qwen 3.8 27B FP8 on dual AMD R9700 (RDNA4) GPUs via vllm-radiance, achieving 5000 tok/s prefill and 130+ tok/s code generation. They're asking whether other architecture-specific inference engines can do better, especially for Qwen 3.8 Flash Next, which vllm-radiance doesn't yet support and llama.cpp performance would be a regression.
More from Infra
- Drax datacentre would burn 4.9m tonnes of wood a year, emissions near double Gatwick flights — nordicinst · 2026-10-06
- Singapore data center operator DayOne files for US IPO after H1 revenue tripled to $512M — zephyr_z9 · 2026-10-06
- Strata claims 6GB VRAM can match RTX 5090-level inference, with full Qwen4 support planned — lxfater · 2026-10-06
- ComfyUI benchmark: Flux 1 dev FP8 tested across 13 GPUs — Ok_Contribution8157 · 2026-10-06
- ODS offers one-click local AI: auto hardware detection, model setup, agents and plugins — Teknium · 2026-10-06
- PlanetScale Engineer Explains Kubernetes Feedback Loops by Running Postgres by Hand — bibryam · 2026-10-06