A plain-C engine runs 230B MiniMax-M2 from disk on a 32GB CPU-only laptop
WritHerAI · reddit · 2026-10-06
- A developer open-sourced Picchio (MIT), a plain-C inference engine for MoE models larger than RAM: dense weights stay in memory while experts stream from SSD on demand, with a cache for hot experts—inspired by Colibri.
- It now supports MiniMax-M2 (230B): 122 GB in INT4, streaming nearly everything from disk.
- On a 12-core laptop with 32 GB RAM, a basic NVMe and no GPU, it hits 0.48 tok/s with a 20 GB expert cache. Slow, but it runs—and cache size turned out to be the only factor that really matters.
- Correctness: forward pass matches MiniMax's original code on a small test model (max logit diff 1e-6).
- Caveats: with 128 GB RAM or a good GPU, llama.cpp is much faster; only tested on Windows with Intel CPUs, and not on M2.5/M3.
More from Infra
- OpenAI to fund Nicholas Nethercote's work speeding up the Rust compiler — charliermarsh · 2026-10-07
- Google signs 20-year nuclear PPA with Constellation: 3,590 MW and $4.3B investment — demian_ai · 2026-10-07
- NVIDIA's Vera CPU spins up 2,000 sandboxes in 27.4s, 8x faster than rivals — mattturck · 2026-10-07
- TablePlus launches a VM built for Apple Silicon and AI Agents with GPU acceleration and MCP — film_girl · 2026-10-07
- Realtime inference startup Reactor raises Series A led by Lightspeed, with Nvidia joining — buckymoore · 2026-10-07
- Terse, an open-source alternative to Cloudflare Durable Objects, joins YC F26 — mertdumenci · 2026-10-07