Qwen3.8 27B hits 280 tok/s on 2x R9700 with MXFP4, beating FP8 at hardware limits
whodoneit1 · reddit · 2026-09-02
- The author kept optimizing dual AMD Radeon R9700 local inference; the community has grown from 5 to 1,200 collaborating developers
- Built MXFP4 support on DeadCode's radiance image using W4A8 kernels, first matching then clearly beating FP8 throughput — now near the hardware limit of these cards
- BetterBench decode: 280 tok/s on JSON, 254 math, 226 code, 148 chat; prefill sustains 3,831 tok/s even at 94k prompt tokens
- KV cache reaches 940k tokens; the full image and repo are open source
More from Infra
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- LeCun urges stronger cybersecurity for Neoclouds to prevent rogue AI takeovers — ylecun · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02
- Asus AI PC Price Jumps 50%, Speculating on Upcoming DGX Spark Hike — mountainyoo · 2026-09-02
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02
- Microsoft Research papers on LLM data infrastructure win awards at VLDB 2026 — jm_alexia · 2026-09-02