colibri: Pure-C Zero-Dependency Engine Streams MoE Experts From Disk to Run Frontier Models Locally
JustVugg · github · 2026-09-10
colibri, trending on GitHub with 27K+ stars, is a tiny pure-C inference engine with zero dependencies that runs frontier MoE models on hardware you already own.
The trick: expert weights are streamed from disk on demand instead of residing in memory, sidestepping VRAM/RAM capacity limits. "Tiny engine, immense model" — a notable entry in the local-inference space.
More from Infra
- ONcompute launches dedicated GPU inference rental with upfront transparent pricing — w1kke · 2026-09-10
- MiniMax H3 local video ecosystem roundup: W4A8 quants, pixel-art guide, faster LoRA training — optimisticalish · 2026-09-10
- Nvidia and Palantir partner to run supply chains with AI, starting with Nvidia's million-part operation — The Decoder · 2026-09-10
- AMD RDNA3 guide: running MiniMax H3 video gen in ComfyUI at ~4 min per 5s clip — Big_Extension_9987 · 2026-09-10
- NVIDIA brings CUDA to Windows on Arm ahead of October RTX Spark AI laptops — kimmonismus · 2026-09-10
- HBM seen topping 40% revenue share next year as Samsung challenges SK Hynix's lead — zephyr_z9 · 2026-09-10