Qwen3.8 Flash Next GGUF benchmark: IQ3_S the sweet spot, 42.7M tokens tested
lxfater · x · 2026-10-07
- Benjamin Marie benchmarked 12 GGUF quantizations of Qwen3.8 Flash Next (Q4 down to Q1, 42.7M generated tokens) against the 354 GB BF16 original.
- Every standard quantization clears a 95% accuracy-recovery threshold; the real differences show up in token efficiency. Below IQ3S, performance on long-horizon agentic coding degrades sharply.
- Notably, quants that scored well on non-agentic evals (like Q1 variants) can fail badly on long-horizon tasks—quantization hurts agentic work far more than standard benchmarks suggest.
- The Coder variant preserves coding accuracy well but sacrifices substantial general knowledge and scientific reasoning.
- lxfater recommends GSQ-RCO IQ3S for local runs: slightly larger than UD Q2 K XL but clearly better in both token efficiency and accuracy.
More from Infra
- Solo dev hits Together AI rate limits running parallel agents, seeks generous API credits — Correct_Positive_108 · 2026-10-07
- Ramjet: an open-source local alternative to NVIDIA Dynamo for multi-GPU inference — DoggoProfessor959 · 2026-10-07
- NVIDIA's NeMo-DCR Cuts Trillion-Parameter RL Weight Sync from 87.5 min to 150s — nvidia · 2026-10-07
- Used PS5 Pro hits $1,399 at GameStop as AI datacenters squeeze memory supply — aakashgupta · 2026-10-07
- Musk: xAI will build and run its Terafab itself, TSMC may only sublease part of it — MickeySteamboat · 2026-10-07
- "72% of the intelligence with 3.8% of the GPUs": Mistral's compute-efficiency ratio sparks debate — cyb3rops · 2026-10-07