Run a 744B MoE model on consumer hardware
Prompt Engineering · youtube · 2026-07-20
A demo of Colibri running GLM 5.2 on consumer hardware with about 25 GB of RAM.
The post highlights two key facts:
- the model is a 744B MoE
- only about 40B parameters are active per token
Because MoE routes only a subset of experts each step, the author argues that only roughly 11 GB of weights change token to token. The takeaway is that extremely large models can sometimes be made usable on modest hardware through routing and careful local execution.
Related event: Open-Source Colibri Runs 744B GLM-5.2 on Consumer Laptops(2 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11