strata-mlx runs a 125B MoE model on a 16GB Mac at 6-7 tok/s via SSD expert streaming
yibie · reddit · 2026-10-10
Strata can run Qwen3.8-Flash-Next, a 125B-parameter MoE model, on a 12GB GPU — but macOS is out of scope. So the author spent a day and a half building strata-mlx, an unofficial MLX engine modeled on Strata's design that reads Strata's GGUF files directly with no weight changes. The trick: keep as many experts in RAM as fit and stream the rest from SSD. Benchmarking by capping the Mac's memory: the 66GB Q20 file runs at 6-7 tok/s on a simulated 16GB machine, 19-23 on 24GB, 28-47 on 32GB, with 1,537-token prompts processed at 273-287 tok/s. A 48GB Mac should hold Q20 entirely in RAM. Code, logs and open problems are on GitHub.
More from Infra
- $500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context — Ok-Breadfruit-3523 · 2026-10-11
- SpaceX reportedly starting Terafab in December: a 100M sq ft chip megafactory — bennash · 2026-10-11
- OpenAI's Jalapeño chip trades kernel difficulty for memory bandwidth, betting on AI-written kernels — nrehiew_ · 2026-10-11
- Google Web AI lead Jason Mayes grew client-side JavaScript AI usage 2500x in 5 years — jason_mayes · 2026-10-11
- Dev in prod: running a newsletter side project on a cheap persistent VM — davidcrawshaw · 2026-10-11
- Zeeg: persistent VMs for agents are wrong, ephemeral sandboxes are the present — zeeg · 2026-10-11