strata-mlx runs a 125B MoE model on a 16GB Mac at 6-7 tok/s via SSD expert streaming

yibie · reddit · 2026-10-10

Strata can run Qwen3.8-Flash-Next, a 125B-parameter MoE model, on a 12GB GPU — but macOS is out of scope. So the author spent a day and a half building strata-mlx, an unofficial MLX engine modeled on Strata's design that reads Strata's GGUF files directly with no weight changes. The trick: keep as many experts in RAM as fit and stream the rest from SSD. Benchmarking by capping the Mac's memory: the 66GB Q20 file runs at 6-7 tok/s on a simulated 16GB machine, 19-23 on 24GB, 28-47 on 32GB, with 1,537-token prompts processed at 273-287 tok/s. A 48GB Mac should hold Q20 entirely in RAM. Code, logs and open problems are on GitHub.

Original post →

More from Infra

Infra channel →