Inco Splash hits 144 tok/s on Qwen3.8-27B M5 Max, 3x faster than Ollama

ResearchCrafty1804 · reddit · 2026-09-19

Open-source inference engine Inco Splash, built for Apple Silicon, runs Qwen3.8-27B at 144 tok/s on an M5 Max — up to 3x Ollama and 2x oMLX, 4x with agent fan-out. One-command setup, works with Claude Code, OpenCode, Codex, and LM Studio.

Related event: Open-Source Inco Splash Engine Runs Qwen3.8-27B at 144 tok/s on M5 Max(3 posts)→

Original post →

More from coding & agent

coding & agent channel →