GLM-4.7 Flash quant benchmark: MXFP4 boosts prompt processing 60%, Q4_K_XL fastest generation

tabletuser_blogspot · reddit · 2026-10-05

A detailed llama.cpp Vulkan benchmark on an AMD Ryzen 7 6800H iGPU (Radeon 680M, 64GB DDR5) compared three 16-17GB 30B.A3B MoE quantizations (averaged over 3 runs):

| Quant | pp512 | tg128 |

|---|---|---|

| MXFP4MOE | 258.36 t/s | 11.66 t/s |

| Q4KM | 218.22 t/s | 12.09 t/s |

| Q4KXL | 160.31 t/s | 13.13 t/s |

Key findings:

Verdict: use Q4KXL for chat/streaming, MXFP4MOE for RAG/long context.

Original post →

More from Infra

Infra channel →