Slipstream runs a 95.5GiB Qwen model on a 64GB Mac at 41-52 tok/s, 1.76x faster than llama.cpp

SnooPredictions515 · reddit · 2026-10-02

Developer npanj released Slipstream, a compiled C++ Metal inference engine for Apple Silicon with native SSD expert streaming and speculative drafting, enabling a 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac.

Results:

How to use:

bash

git clone https://github.com/npanj/slipstream.git && cd slipstream && make -j4

sudo sysctl iogpu.wiredlimitmb=59392 required on 64GB Macs

./slipstream serve --model /models/qwen38-flash-next-v3 --port 8090

Open source on GitHub; weights on Hugging Face.

Original post →

More from Infra

Infra channel →