Dual CMP 170HX (128GB HBM) runs GLM-5.3-Flash at 384K context, ~90 tok/s with EXL3
Prudent_Appearance71 · reddit · 2026-10-08
A r/LocalLLaMA user shares a stable local setup for GLM-5.3-Flash (320B MoE, 18B active) on two 64GB CMP 170HX cards:
- Setup: EXL3 3.05bpw quant, 125GB weights fully resident in HBM (HBM-first approach, no expert streaming), DFlash2 speculative decoding, Q8 KV cache, 384K usable context (393K request budget), 90 tok/s.
- Quant quality: EXL3 uses trellis-based mixed precision — experts compressed hard while sensitive paths keep more bits (lmhead at 6-bit). Published 3.0bpw fidelity results (51,175 held-out positions) show 93% top-1 agreement and 0.050 mean KLD, beating UD-IQ4XS (88.18%) at only 125GB.
- Comparison: same machine also ran Qwen3.8-Flash-Next (AWQ INT4 + FP8 on vLLM); author notes engines and speculative decoding differ, so it's not an apples-to-apples benchmark.
Repo and demos: glm53-flash-cmp170hx-exl3 on GitHub.
More from Infra
- Reddit asks: have AI scaling laws hit their limit, or is the compute buildout just starting? — StupidDialUp · 2026-10-08
- TRIAGE stabilizes native NVFP4 RL training, hits full-precision quality at 2.3x throughput — InfiX-ai · 2026-10-08
- WSL Containers Now Generally Available: Run Linux Containers Natively on Windows — pavandavuluri · 2026-10-08
- ai& Says It's Japan's Largest Dedicated Inference Provider, Teases Post-Training Offerings — DavidBennett__ · 2026-10-08
- Microsoft unveils Surface Laptop Ultra with Nvidia RTX Spark SoC from $2,599 — Ars Technica AI · 2026-10-08
- DeepSeek V4.1 shrinks cache 437x, Flash beats V4-Pro 39 vs 36 at half the cost — DeepLearningAI · 2026-10-08