Running a 744B MoE on a 25GB laptop: Colibri streams experts from disk like a weight JIT
thisdudelikesAI · x · 2026-10-07
Colibri demonstrates running a 744B-parameter MoE model on a laptop with just 25GB of RAM. The core idea: don't load the model, place it.
- Only 40B parameters fire per token, with 11GB of weights changing between tokens, so the full model never needs to fit in fast memory
- The dense parts (attention, shared experts, embeddings) stay in RAM at int4 (9.9GB)
- The 19,456 routed experts (19MB each, 370GB total) live on disk and stream in only when the router picks them
- A lookahead thread prefetches the next layer's experts — routing is 71.6% predictable one layer ahead
- It logs which experts your workload uses and pins the hottest ones, getting faster over time
The author calls it a JIT for weights — a notable path to running huge models on consumer hardware.
More from Infra
- llama.cpp merges GLM5Next MTP support — GLM 5 Flash now runs locally — jacek2023 · 2026-10-07
- Europe's AI gap is compute: infrastructure moves in years while AI moves in months — ingliguori · 2026-10-07
- Chrome 155 ships JPEG XL: 30-50% better compression, Rust decoder for memory safety — jedisct1 · 2026-10-07
- Dual 7900XTX Owner Weighs Waiting for LPDDR6 Platforms vs Going WRX80 — MikeSouto · 2026-10-07
- POC 2026 talk shows container escape through the NVIDIA GPU driver despite read-only limits — evilsocket · 2026-10-07
- Google Releases SAM, a P2P Network Letting AI Agents Discover and Call Each Other's Tools — thisguyknowsai · 2026-10-07