Qwen 27B on a single R9700: 3-bit rotation quant buys 569K-token cache and 12x faster agent turns

evp-cloud · reddit · 2026-10-04

A developer detailed how to run Qwen3.8 27B locally on one AMD Radeon AI PRO R9700 (32GB, 300W) with speculative decoding and a carefully scoped 3-bit quantization:

Weights, calibration sources and hashes are published on Hugging Face; the 3-bit add-on is only 9.55GB.

Original post →

More from Infra

Infra channel →