Qwen3.8-27B quant shootout: MXFP4 tops math benchmarks while Q6_K_L wins code
JurBank · reddit · 2026-10-12
A Reddit user benchmarked three Qwen3.8-27B quants on an R9700 using gsm8k and humaneval inside Docker.
Key findings:
- Swift-1.5 MXFP4 (out Q6K): best math (0.96 gsm8k at xhigh thinking) and fastest (3.6-6.1 min), but worst code (0.62-0.67 humaneval)
- Swift Q6K: 0.76 code, yet slowest overall (9.8-18.2 min)
- UD-Q6KL: best code (0.78-0.80), worst math (0.88-0.89)
The hypothesis that Q6KL would dominate accuracy and MXFP4 would trail was largely overturned. The poster asks why MXFP4 excels at math but lags on code. Sample sizes (50-100 problems) are small, so results need larger-scale replication.
More from Infra
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12
- VitalOps launches agentic inference optimization, 2.6x median speedup — abhijithneil · 2026-10-12
- Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s — human_in_the_looop · 2026-10-12
- Siemens Brings Nvidia Omniverse into Digital Twin Composer to Pave the Way for Physical AI — RevLebaredian · 2026-10-12
- OpenRouter token traffic explodes from 2T to 379T/month, open-weight models at 75% — Beth_Kindig · 2026-10-12