Standard evals can't tell local quants apart, but a self-built agentic bench can (77% vs 98%)

FeydRowan · reddit · 2026-10-11

The author spent weeks measuring quant quality in local setups. Five quants of Qwen3.8-27B (UD-Q4KXL through Q80, plus NVFP4) score identically on GSM8K, MMLU-Pro and IFEval — within ±2 pts — even though Q4 picks a different token than Q8 once every 28. Differences aren't even ordered: the least faithful quant won IFEval. Only KLD/top-1 agreement ranks quants, and it takes 40 seconds. Q4KXL stays the daily driver: +33% decode speed and 7.5× context vs Q80.

Across models, GSM8K is saturated (96.8–98.5% for Qwen 27B, Flash-Next, Gemma 4 26B/31B, Claude Sonnet 5.5). The gap appears only on agentic coding tasks the models can't have seen: on a self-built bench (bugs injected into the author's own repos, hidden tests, 112 runs per model via OpenCode), 27B NVFP4 solved 77% vs Flash-Next NVFP4's 98% — 1/12 vs 12/12 on the six hardest tasks; SWE-rebench corroborates (69% vs 82%, p≈0.007). Flash-Next at 3 bits (GSQ IQ3S on Strata) locally matches NVFP4 on 11 of 12 discriminating tasks but misses one specific bug every time. Harness matters too: OpenCode was 2–4× faster and passed every safety probe. Setup: dual RTX 5070 Ti, 96GB DDR5, llama.cpp for GGUFs, vLLM TP=2 for NVFP4.

Original post →

More from Models

Models channel →