Inco Splash hits 144 tok/s running Qwen3.8-27B on an M5 Max, 3x faster than Ollama
TheMoonMidas · x · 2026-09-19
Inco released Splash, an open-source inference engine built specifically around Apple silicon, with the DFlash 2 draft model pre-integrated and runnable in two commands.
- Qwen3.8-27B now runs at 144 tok/s on an M5 Max MacBook Pro, up from 70 tok/s with DFlash 2 a month ago
- Claims up to 3x the decode speed of Ollama and 2x oMLX, nearly 4x when an agent fans out into sub-agents
- The engine is co-designed with the model rather than being a generic runtime
Related event: Inco Splash Open-Source Engine Runs Qwen3.8-27B at 144 tok/s on M5 Max(2 posts)→
More from Infra
- Jev scales on Modal, one of only 4 subprocessors listed, to meet surging demand — AAAzzam · 2026-09-19
- Tuning Qwen3 27B as a coding agent on 2x3090s cuts turn latency from 28s to 7s — bolts98 · 2026-09-19
- AgentZip: memory compression for parallel agent sandboxes cuts memory up to 8.7x — rohanpaul_ai · 2026-09-19
- US products quietly build on Chinese open-weight models as one firm cuts spend by ~100x — generativist · 2026-09-19
- Turbines sold out through 2030: 25 categories of datacenter power gear face multi-year backlogs — CatAstro_Piyush · 2026-09-19
- Why OpenAI and Anthropic are suddenly buying tiny 20-30MW data centers — abhiadesai · 2026-09-19