Qwen 27B on a single R9700: 3-bit rotation quant buys 569K-token cache and 12x faster agent turns
evp-cloud · reddit · 2026-10-04
A developer detailed how to run Qwen3.8 27B locally on one AMD Radeon AI PRO R9700 (32GB, 300W) with speculative decoding and a carefully scoped 3-bit quantization:
- Not a blanket 3-bit quant: only the large projection matrices (MLP, attention, Gated DeltaNet; 24.3B of 27B params) are quantized to 3.1 bits/weight with per-128-block scales, after a rotation that spreads outliers, GPTQ-calibrated on 293K tokens. Embeddings, output head and norms stay MXFP4; generation uses FP8 activations (W3A4).
- Gains: +15% decode speed (179.7 tok/s), 2.25x KV cache, and a reusable prefix cache of 569,878 tokens in coding mode — each request gets 262,144-token context, two full-length requests fit at once.
- Agent impact: coding agents resend the whole conversation each turn; now only new tokens are read — later turns start up to 12x sooner, full sessions finish 6x faster, and re-reading a 258K-token document drops from 134s to 2.7s.
- Accuracy cost: GSM8K and HumanEval within noise; MMLU-Pro drops 2-3 points. A one-command MXFP4 mode remains available from the same download for accuracy-first users.
Weights, calibration sources and hashes are published on Hugging Face; the 3-bit add-on is only 9.55GB.
More from Infra
- Rural Queensland faces a $31B Anthropic datacentre, and locals aren't happy — nordicinst · 2026-10-04
- America's data center fight previews what's coming for the rest of the world — pstAsiatech · 2026-10-04
- US-China RISC-V collaboration deepens as open chip architecture moves into data centers — pstAsiatech · 2026-10-04
- Thoughtworks Engineer Spends 4 Weeks Testing Whether Local Models Are Viable for Coding — bibryam · 2026-10-04
- Cloudflare's open-weights decision model Clef scores all answers in one pass, 53ms per decision on RTX 5090 — michellechen · 2026-10-04
- Tencent to lease ~100k chips from Oracle for ~$7bn; BoE flags agent risk — MirelaXhota · 2026-10-04