Inco Splash open-source engine hits 144 tok/s on Qwen3.8-27B in an M5 Max
Lianhuiq · x · 2026-09-19
Inco released Inco Splash, an open-source inference engine built around the model and Apple silicon: Qwen3.8-27B runs at 144 tok/s on an M5 Max MacBook Pro. Claimed decode speed is up to 3x Ollama, 2x oMLX, and nearly 4x when an agent fans out into sub-agents.
Related event: Inco Splash Open-Source Engine Runs Qwen3.8-27B at 144 tok/s on M5 Max(2 posts)→
More from Infra
- CoreAI Zoo hits App Store: run Qwen, Gemma and more fully on-device on iPhone — pcuenq · 2026-09-19
- MiniMax H3 ecosystem roundup: quantization, multi-GPU inference, LoRA tooling and workflows in one week — Just_Lingonberry_352 · 2026-09-19
- SpaceX to make gas-turbine blades in-house as AI power crunch strains supply chain — AccBalanced · 2026-09-19
- Running the 27B Bansai local model on just 5GB of RAM — hands-on test begins — draginol · 2026-09-19
- Inco Splash hits 144 tok/s running Qwen3.8-27B on an M5 Max, 3x faster than Ollama — TheMoonMidas · 2026-09-19
- Looped transformers study: 7.4B growth model matches GPT-3 13B with 20x less compute — burny_tech · 2026-09-19