COLIBRI Runs 744B GLM-5.2 Model on 25GB RAM Without GPU
COLIBRI demonstrated a disk streaming hack to run the massive 744B parameter GLM-5.2 model on a 25GB RAM machine without a GPU. Although the inference speed is extremely slow, it proves that consumer-grade hardware can run cutting-edge MoE models.
2026-07-13 ~ 2026-07-14 · 3 related posts
- Running GLM-5.2 on a 25GB Machine — paulabartabajo_ · 2026-07-13
- 744B Model Runs on 25GB Machine — alexcovo_eth · 2026-07-13
- Disk-based streaming inference solution for GLM-5.2 — karminski3 · 2026-07-14