GLM-4.7 Flash quant benchmark: MXFP4 boosts prompt processing 60%, Q4_K_XL fastest generation
tabletuser_blogspot · reddit · 2026-10-05
A detailed llama.cpp Vulkan benchmark on an AMD Ryzen 7 6800H iGPU (Radeon 680M, 64GB DDR5) compared three 16-17GB 30B.A3B MoE quantizations (averaged over 3 runs):
| Quant | pp512 | tg128 |
|---|---|---|
| MXFP4MOE | 258.36 t/s | 11.66 t/s |
| Q4KM | 218.22 t/s | 12.09 t/s |
| Q4KXL | 160.31 t/s | 13.13 t/s |
Key findings:
- The iGPU lacks native FP4 (fp4: 0), so MXFP4 is emulated at runtime, yet MoE structure + extreme quantization still cut prompt-processing compute/memory reads, giving 60% faster PP with Flash Attention
- Generation is purely memory-bandwidth bound (50-65 GB/s); Q4KXL's layout prefetches better on RADV, 13% faster than Q4KM
- Very tight variance (±0.02-0.06 t/s); Q4KXL's first-run PP outlier was cold cache
Verdict: use Q4KXL for chat/streaming, MXFP4MOE for RAG/long context.
More from Infra
- 539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30% — HankYeomans · 2026-10-05
- GLM 5.3 flash on dual DGX Sparks gets 50-90% decode boost with new open recipe — swiebertjee · 2026-10-05
- Qualcomm's Snapdragon to power next-gen AI assistants for Meta and OpenAI — ryanshrout · 2026-10-05
- Spite: a modular Rust inference engine where every model, GPU and op is pluggable — giveen · 2026-10-05
- Only 5% of chips tape out right first time — repeat spins, not fabs, may bottleneck custom silicon — ai · 2026-10-05
- Nvidia launches Open Agent Safety Platform to rein in rogue AI agents — dl_weekly · 2026-10-05