Local Qwen 3.8 Flash Next benchmark jumps from 91 to 98/100 with AWQ quants and patched vLLM
WonderRico · reddit · 2026-09-05
Blogger WonderRico updated his local LLM benchmark recipe for Qwen 3.8 Flash Next, raising the score from 91 to 98/100.
- New stack: AWQ W4A16 weights plus INT4 n-gram PLE quant (both on Hugging Face), served via vLLM patched with files from the primitive-ai repo
- Goal: load both weights and PLE in 4-bit on sm89/sm120 GPUs; slower than the previous patched SGLang setup but a large quality jump
- Hit 98/100 at both medium and xhigh reasoning — best score and most efficient combo on his bench
- He'll investigate whether the delta comes from the engine patches or the quants; graphs and data are public
More from Infra
- Agent bottleneck wasn't latency: token-per-minute ceiling capped product at 2.2 turns/min — Initial_Orange2985 · 2026-09-05
- Back-of-Envelope: Global Compute Equals ~20M H100s; Inference Dominates, Pretraining Far Less Efficient Than Biology — JosephJacks_ · 2026-09-05
- Tenstorrent and aiand Launch JapanFold, Free Open-Source Drug Discovery Models on Galaxy Clusters — DavidBennett__ · 2026-09-05
- Aer Lingus says 50% of long-haul fleet now has Starlink, full coverage by year-end — elonmusk · 2026-09-05
- Running Qwen3.8-27b-mlx Locally on an M5 MacBook to Update a Website and Generate Assets — walkingriver · 2026-09-05
- Agent Substrate brings instant suspend/resume and 10x density to K8s AI agents — davemccollough · 2026-09-05