GLM-4.7 quant showdown on a Radeon 680M iGPU: MXFP4 vs Q4_K_XL vs Q4_K_M

tabletuser_blogspot · reddit · 2026-08-19

Benchmarks of GLM-4.7-Flash quant variants (GGUF header reports deepseek2 30B.A3B MoE) on an Acemagic mini-PC (Ryzen 7 6800H + Radeon 680M iGPU, 64GB DDR5, Ubuntu + llama.cpp Vulkan): MXFP4 MoE (15.79GiB), UD-Q4KXL (16.31GiB), Q4KM (17.05GiB).

Results (3-run averages, Flash Attention on):

| Format | Prompt processing pp512 | Generation tg128 |

|---|---|---|

| MXFP4 MoE | 258.36 t/s | 11.66 t/s |

| Q4KM | 218.22 t/s | 12.09 t/s |

| UD-Q4KXL | 160.31 t/s | 13.13 t/s |

Analysis: models exceed iGPU memory, so performance is bound by DDR5 bandwidth (50–65GB/s); the platform lacks native FP4 (fp4:0), yet MXFP4 still wins prompt processing thanks to MoE's small active compute. Generation is purely bandwidth-bound, where Q4KXL's layout suits RADV driver prefetching best. Standard deviations are tiny across runs.

Recommendations: pick Q4KXL for chat/streaming; MXFP4MoE for RAG/long context with its 60% prompt-processing advantage.

Original post →

More from Infra

Infra channel →