Slipstream runs a 95.5GiB Qwen model on a 64GB Mac at 41-52 tok/s, 1.76x faster than llama.cpp
SnooPredictions515 · reddit · 2026-10-02
Developer npanj released Slipstream, a compiled C++ Metal inference engine for Apple Silicon with native SSD expert streaming and speculative drafting, enabling a 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac.
Results:
- Same V3 checkpoint (34k+ downloads on HF): llama.cpp fork manages 23.1 tok/s; Slipstream hits 41–52 tok/s, a 1.76x average speedup
- No long-context collapse: across 3,086 live coding-session requests, speed stays flat at 33–44 tok/s all the way to 130,000 tokens
- Benchmarks across GSM8K, MATH-500, logic, Python/Rust coding and tech writing show 1.65–1.82x gains; TTFT drops from 1,863ms to 1,290ms
How to use:
bash
git clone https://github.com/npanj/slipstream.git && cd slipstream && make -j4
sudo sysctl iogpu.wiredlimitmb=59392 required on 64GB Macs
./slipstream serve --model /models/qwen38-flash-next-v3 --port 8090
- Standard OpenAI-compatible API works with Claude Code, OpenCode, curl, etc.
- First launch pre-processes multi-shard GGUF in 5–7 minutes; subsequent loads take 10–15s
- Optional Swift KV-Sparse model variant also available
Open source on GitHub; weights on Hugging Face.
More from Infra
- MachGen pushes MiniMax H3 past its 15s cap with 30-second continuous video — MiniMax_AI · 2026-10-02
- Redditor builds fully local LLM-powered radio site on two DGX Sparks and a 5090 — jwhh91 · 2026-10-02
- VC quip: many neoclouds are closer to 95% than five nines of reliability — saranormous · 2026-10-02
- Report: lenders demand up to 25% collateral from Nvidia as GPU-backed loans wobble — GaryMarcus · 2026-10-02
- Microsoft Backs Snowflake-Led Effort to Standardize Business Metrics for AI — xiaosun86 · 2026-10-02
- GLM-5.3-Flash NVFP4 benchmarks show no per-user speedup beyond 8 concurrent requests — TheZachMueller · 2026-10-02