Qwen3.8-27B on AMD R9700 hits 227 tok/s with lossless block-diffusion drafter
samsja19 · x · 2026-08-27
Lucebox engine now serves Qwen3.8-27B on a single AMD Radeon AI PRO R9700 (32GB, RDNA4) using DFlash2 block-diffusion drafter, achieving up to 227 tok/s on code, 208 tok/s on HumanEval, and 133 tok/s on math. Using Unsloth's UD-IQ4XS quant, it matches 8-bit reference with KL divergence 0.018 and 94% top-1 agreement, and is lossless.
More from Infra
- 4D Gaussian Splatting runtime cut to 5 mins, memory needs halved — janusch_patas · 2026-08-27
- ETH Zurich undergrads tape out full ASIC chips with full funding — AccBalanced · 2026-08-27
- Figure's Index: 43k weekly users feeding robot data, $15M paid, $1B committed — coreylynch · 2026-08-27
- OpenAI reveals AI-assisted chip design, accelerating development of first-gen chip — AccBalanced · 2026-08-27
- Node for truly free model and node cache in Stable Diffusion — JustLookingForNothin · 2026-08-27
- Redis creator Antirez commits to non-profit local inference engine — antirez · 2026-08-27