GLM-5.3-Flash runs 3.3x faster locally with Unsloth's 1-bit GGUFs on 100GB RAM
danielhanchen · x · 2026-09-04
Unsloth published a full guide to running GLM-5.3-Flash (Z.ai's ox-alpha, a 320B-parameter/18B-active multimodal open model) locally, with 1.6–3.4x faster inference via MTP and optimized long-context decoding.
- Trained on 30T tokens with hybrid sparse+linear attention; claimed to rival Claude Opus 4.8 on coding and agentic benchmarks
- Dynamic 1-bit GGUF: 93GB vs 642GB BF16, retaining 71% of top-1% accuracy; 3-bit is 76% smaller with 87% accuracy
- Hardware: 100GB RAM for 1-bit, 128GB setups (Mac, DGX Spark) for 3-bit
- Three thinking modes (Low/High/Max); DeepSWE tasks suggest temperature=0.95
Runs via Unsloth Desktop or llama.cpp; GGUFs on Hugging Face.
More from Infra
- Kafka exactly-once delivery demystified: the 3-layer design and the catch engineers miss — arpit_bhayani · 2026-09-04
- AMD's Threadripper Halo Station packs 96 cores and 576GB of HBM3E — ccerrato147 · 2026-09-04
- Pennsylvania voters unite against data centres: 'People are going to get screwed' — 1vuio0pswjnm7 · 2026-09-04
- AeroJEPA fluid foundation model joins NVIDIA's PhysicsNeMo ecosystem — ricardovinuesa · 2026-09-04
- Building a €2-2.5k local AI rig for legal RAG and agentic coding: hardware picks debated — whatyathinkk · 2026-09-04
- Dual 3090 owners debate adding more cards: bigger local models vs parallel instances — Blues520 · 2026-09-04