GLM-5.3-Flash quantized to 2.0bpw runs on a single DGX Spark at 18-25 tok/s with vision

pbaylies · x · 2026-09-08

Developer 0xSero released an EXL3 quantization (TR3, 2.0bpw, no pruning) of GLM-5.3-Flash sized for a single NVIDIA DGX Spark. It's validated working at 18-25 tok/s with vision enabled, weights are up on Hugging Face with full chat template config, and the author is soliciting community testing feedback.

Original post →

More from Infra

Infra channel →