Dual CMP 170HX (128GB HBM) runs GLM-5.3-Flash at 384K context, ~90 tok/s with EXL3

Prudent_Appearance71 · reddit · 2026-10-08

A r/LocalLLaMA user shares a stable local setup for GLM-5.3-Flash (320B MoE, 18B active) on two 64GB CMP 170HX cards:

Repo and demos: glm53-flash-cmp170hx-exl3 on GitHub.

Original post →

More from Infra

Infra channel →