A 36GB MacBook can run 35B–480B coding models by streaming experts from SSD
Illustrious-Cup-5895 · reddit · 2026-07-26
A developer says they can run 35B–480B coding models on a 36 GB MacBook by streaming MoE experts from SSD instead of loading everything into RAM. The app ships as a self-contained macOS .dmg, supports localhost integration with tools like Cursor and Cline, and includes benchmarks that highlight both wins and failed experiments.
Key results:
- Qwen3.6-35B-A3B reaches about 13 tok/s at a 10 GiB cache and 19 tok/s at 14 GiB.
- Laguna 118B-A8B runs at about 2.8 tok/s, good enough for batch use.
- Moving the streamed expert file to faster storage produced a 2.7× gain.
- Prefetching, dual-SSD striping on a shared bus, and hot-expert reservation all underperformed.
More from coding & agent
- Anthropic’s Claude agent lead says delegation, not prompting, is now the bottleneck — petergyang · 2026-07-26
- Nine Claude habits that can help users avoid hitting usage limits — Roger_M_Taylor · 2026-07-26
- Anthropic says Claude Code now lands 65% of product engineering PRs — Roger_M_Taylor · 2026-07-26
- Open-source book-to-skill converts entire books into chapter-by-chapter Claude skills — Roger_M_Taylor · 2026-07-26
- Multiplayer multi-agent systems turn agent naming and scope into team-wide decisions — edgarpavlovsky · 2026-07-26
- Six AI windows coding at once now looks like a normal workday to this developer — gandamu_ml · 2026-07-26