Qwen Flash Next IQ4_XS beats 27B FP8 on MMLU-Pro, GPQA and GSM8K in community eval
smallDeltaBigEffect · reddit · 2026-09-25
A community eval (run via Codex) comparing Qwen 3.8 Flash Next IQ4XS (vllm + r9v) against Qwen 3.8 27B FP8 (vllm + radiance) shows the low-bit quant winning across the board:
| Benchmark | Flash-Next IQ4XS | 27B FP8 | Winner |
|---|---|---|---|
| MMLU-Pro | 83.8% | 75.0% | Flash-Next |
| GPQA Diamond | 42.5% | 27.5% | Flash-Next |
| GSM8K | 97.5% | 90.0% | Flash-Next |
| German MGSM | 92.5% | 90.0% | Flash-Next |
| IFEval | 89.6% | 89.6% | Tie |
The cost is speed: FP8 27B finished the 250-case "fast" trial in 36m 27s vs 2h 05m 47s for Flash-Next IQ4XS. Hardware: 64GB DDR5 + 2×R9700 (recent r9v updates not included). With 128GB DDR5, tcclaviger's mxfp4 quant could roughly double Flash-Next throughput.
The author pushes back on the recurring claim that FP8 27B beats low quants of QFN — IQ4XS already counts as "low," yet neither this eval nor daily use supports it.
More from Infra
- Nutanix Acquires Ryax to Squeeze More From Idle GPUs, Cutting Node-Hours 62% in Tests — shashib · 2026-09-25
- Former Intel CEO calls HBM "lousy" at Hot Chips 2026 as High Bandwidth Flash looms — Glittering_Depth_722 · 2026-09-25
- MLPerf Training v6.1 adds first LLM post-training benchmark: agentic RL on a 397B model — TheKanter · 2026-09-25
- kvcached brings virtual memory to LLM KV cache, deployed on 10K+ GPUs — techNmak · 2026-09-25
- IEEE plenary talk: micro-optimizations across the full stack, from silicon to models — fooobar · 2026-09-25
- Dev builds local AI GTM workflow, argues the next platform entry point is hardware-bound — dotey · 2026-09-25