744B Model Runs on 25GB Machine
alexcovo_eth · x · 2026-07-13
COLIBRI demonstrates a prototype that runs GLM-5.2 (744B parameters) on a consumer-grade machine with 25GB of RAM and no GPU.
The core idea is that the model doesn't need to keep all parameters resident in memory at once. Instead, it keeps a small portion in RAM and streams the rest from disk on-demand. While disk speed bottlenecks generation speed, it proves that running giant models doesn't strictly require massive VRAM. The project is open-source under Apache-2.0 and has garnered around 2.1k stars.
Related event: COLIBRI Runs 744B GLM-5.2 Model on 25GB RAM Without GPU(3 posts)→
More from Infra
- Vercel AI Gateway data shows Anthropic, OpenAI and Google at 97.09% spend share — cramforce · 2026-07-21
- NVIDIA starts rolling out 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-21
- Mustafa Suleyman says Microsoft is preparing for an OpenAI exit, while a new chip costs 30% less than GB200 — thoefler · 2026-07-21
- Microsoft and Mistral sign multi-billion-dollar deal to expand AI infrastructure in Europe — The Decoder · 2026-07-21
- Speculative decoding boosts Qwen3.6-27B on one 5090, but slows crowded servers — luke_pacman · 2026-07-21
- NVIDIA says Blackwell Ultra hit 1,648 TFLOPs per GPU on DeepSeek-V3 671B training — NVIDIAAI · 2026-07-21