Qwen Flash Next IQ4_XS beats 27B FP8 on MMLU-Pro, GPQA and GSM8K in community eval

smallDeltaBigEffect · reddit · 2026-09-25

A community eval (run via Codex) comparing Qwen 3.8 Flash Next IQ4XS (vllm + r9v) against Qwen 3.8 27B FP8 (vllm + radiance) shows the low-bit quant winning across the board:

| Benchmark | Flash-Next IQ4XS | 27B FP8 | Winner |

|---|---|---|---|

| MMLU-Pro | 83.8% | 75.0% | Flash-Next |

| GPQA Diamond | 42.5% | 27.5% | Flash-Next |

| GSM8K | 97.5% | 90.0% | Flash-Next |

| German MGSM | 92.5% | 90.0% | Flash-Next |

| IFEval | 89.6% | 89.6% | Tie |

The cost is speed: FP8 27B finished the 250-case "fast" trial in 36m 27s vs 2h 05m 47s for Flash-Next IQ4XS. Hardware: 64GB DDR5 + 2×R9700 (recent r9v updates not included). With 128GB DDR5, tcclaviger's mxfp4 quant could roughly double Flash-Next throughput.

The author pushes back on the recurring claim that FP8 27B beats low quants of QFN — IQ4XS already counts as "low," yet neither this eval nor daily use supports it.

Original post →

More from Infra

Infra channel →