GLM-5.3-Flash quantized to 2.0bpw runs on a single DGX Spark at 18-25 tok/s with vision
pbaylies · x · 2026-09-08
Developer 0xSero released an EXL3 quantization (TR3, 2.0bpw, no pruning) of GLM-5.3-Flash sized for a single NVIDIA DGX Spark. It's validated working at 18-25 tok/s with vision enabled, weights are up on Hugging Face with full chat template config, and the author is soliciting community testing feedback.
More from Infra
- Running MiniMax H3 video gen on 4GB VRAM: blurry hands at 512p or 6+ hours per clip at native res — unbenannt1 · 2026-09-08
- Celesto AI launches CelestoFS, petabyte-scale durable workspaces for AI agent sandboxes — aniketmaurya · 2026-09-08
- Nvidia's stakes in its own chip buyers hit $99B, up from $7B a year ago — GaryMarcus · 2026-09-08
- Pairing a decade-old RX 480 with RX 7900 GRE boosts llama.cpp MoE inference 36% over single GPU — tabletuser_blogspot · 2026-09-08
- Extrapolating OpenAI's plots: agents may eat 25% of compute within 9-12 months — joshua_clymer · 2026-09-08
- User runs Pinokio with Wan2GP, MiniMax H3, Viggle-Animate on DGX Spark — cocktailpeanut · 2026-09-08