Standard evals can't tell local quants apart, but a self-built agentic bench can (77% vs 98%)
FeydRowan · reddit · 2026-10-11
The author spent weeks measuring quant quality in local setups. Five quants of Qwen3.8-27B (UD-Q4KXL through Q80, plus NVFP4) score identically on GSM8K, MMLU-Pro and IFEval — within ±2 pts — even though Q4 picks a different token than Q8 once every 28. Differences aren't even ordered: the least faithful quant won IFEval. Only KLD/top-1 agreement ranks quants, and it takes 40 seconds. Q4KXL stays the daily driver: +33% decode speed and 7.5× context vs Q80.
Across models, GSM8K is saturated (96.8–98.5% for Qwen 27B, Flash-Next, Gemma 4 26B/31B, Claude Sonnet 5.5). The gap appears only on agentic coding tasks the models can't have seen: on a self-built bench (bugs injected into the author's own repos, hidden tests, 112 runs per model via OpenCode), 27B NVFP4 solved 77% vs Flash-Next NVFP4's 98% — 1/12 vs 12/12 on the six hardest tasks; SWE-rebench corroborates (69% vs 82%, p≈0.007). Flash-Next at 3 bits (GSQ IQ3S on Strata) locally matches NVFP4 on 11 of 12 discriminating tasks but misses one specific bug every time. Harness matters too: OpenCode was 2–4× faster and passed every safety probe. Setup: dual RTX 5070 Ti, 96GB DDR5, llama.cpp for GGUFs, vLLM TP=2 for NVFP4.
More from Models
- Meta's Muse growth slowing: daily active user gains down 62.4% from September surge — AccBalanced · 2026-10-11
- Researcher claims OpenAI exploits user prompts, cites Tao atop a list — basedjensen · 2026-10-11
- 'Nothing new': researcher says engram overfitting is plain overfitting tied to over-parametrization — teortaxesTex · 2026-10-11
- Leaked: OpenAI's unreleased model solved most math problems in a single prompt, ~3h each — Puzzleheaded-King584 · 2026-10-11
- Turing Post's 2026 guide: which LLM benchmarks to use for reasoning, coding, math and agents — TheTuringPost · 2026-10-11
- Musk says RAM supply is the problem as quantized 25GB MoE local run speculated for 2028-29 — elonmusk · 2026-10-11