Inco Splash open-source engine hits 144 tok/s on Qwen3.8-27B in an M5 Max

Lianhuiq · x · 2026-09-19

Inco released Inco Splash, an open-source inference engine built around the model and Apple silicon: Qwen3.8-27B runs at 144 tok/s on an M5 Max MacBook Pro. Claimed decode speed is up to 3x Ollama, 2x oMLX, and nearly 4x when an agent fans out into sub-agents.

Related event: Inco Splash Open-Source Engine Runs Qwen3.8-27B at 144 tok/s on M5 Max(2 posts)→

Original post →

More from Infra

Infra channel →