Open-source Colibri runs GLM-5.2 744B MoE on a 25GB RAM laptop
socialwithaayan · x · 2026-07-21
- An open-source pure-C runtime called colibri claims it can run GLM-5.2, a 744B MoE model, on a machine with about 25GB of RAM.
- The key idea is that only about 40B parameters are active per token, while experts are streamed from disk and a learning cache keeps the hot set resident.
- The post shows measured speeds ranging from 0.05 tok/s on a 25GB dev box to 1.8 tok/s on a 128GB CPU desktop and 5.8–6.8 tok/s on a 6× RTX 5090 setup.
- The runtime is described as single-file, zero-dependency C, with token-exact validation against a full-precision oracle and a demo that reports 9.9GB resident memory after startup.
Related event: Open-Source Colibri Runs 744B GLM-5.2 on Consumer Laptops(2 posts)→
More from Infra
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22