GLM-4.7 quant showdown on a Radeon 680M iGPU: MXFP4 vs Q4_K_XL vs Q4_K_M
tabletuser_blogspot · reddit · 2026-08-19
Benchmarks of GLM-4.7-Flash quant variants (GGUF header reports deepseek2 30B.A3B MoE) on an Acemagic mini-PC (Ryzen 7 6800H + Radeon 680M iGPU, 64GB DDR5, Ubuntu + llama.cpp Vulkan): MXFP4 MoE (15.79GiB), UD-Q4KXL (16.31GiB), Q4KM (17.05GiB).
Results (3-run averages, Flash Attention on):
| Format | Prompt processing pp512 | Generation tg128 |
|---|---|---|
| MXFP4 MoE | 258.36 t/s | 11.66 t/s |
| Q4KM | 218.22 t/s | 12.09 t/s |
| UD-Q4KXL | 160.31 t/s | 13.13 t/s |
Analysis: models exceed iGPU memory, so performance is bound by DDR5 bandwidth (50–65GB/s); the platform lacks native FP4 (fp4:0), yet MXFP4 still wins prompt processing thanks to MoE's small active compute. Generation is purely bandwidth-bound, where Q4KXL's layout suits RADV driver prefetching best. Standard deviations are tiny across runs.
Recommendations: pick Q4KXL for chat/streaming; MXFP4MoE for RAG/long context with its 60% prompt-processing advantage.
More from Infra
- Marvell issues Google warrant tied to custom chip revenue through FY2033 — zephyr_z9 · 2026-08-19
- Weaviate scales test-time compute in search, boosting nDCG to 57.5 — lateinteraction · 2026-08-19
- SALT: CELF-Based Sentence-Level Compression for KV Cache Retrieval — No_Sky9786 · 2026-08-19
- Beyond human intuition: AI designs chip components 500x smaller than engineering limits — ChuckDBrooks · 2026-08-19
- ComfyUI becomes unusable overnight with MiniMax H3, causing system freezes — Fit-Association-448 · 2026-08-19
- Suggestion: Move Anthropic bio AI to Tenstorrent to cut costs — DavidBennett__ · 2026-08-19