Inco Splash hits 144 tok/s on Qwen3.8-27B M5 Max, 3x faster than Ollama
ResearchCrafty1804 · reddit · 2026-09-19
Open-source inference engine Inco Splash, built for Apple Silicon, runs Qwen3.8-27B at 144 tok/s on an M5 Max — up to 3x Ollama and 2x oMLX, 4x with agent fan-out. One-command setup, works with Claude Code, OpenCode, Codex, and LM Studio.
Related event: Open-Source Inco Splash Engine Runs Qwen3.8-27B at 144 tok/s on M5 Max(3 posts)→
More from coding & agent
- Vibe42: terminal-first remote workspace to run coding agents from your phone — yihui_indie · 2026-09-19
- wx-cli + Codex: automating lead capture, CRM, delivery and invoicing via WeChat — lxfater · 2026-09-19
- Roboclaw joins team Discord as a live assistant that reports on all agent sessions — steipete · 2026-09-19
- Scale AI's Alexandr Wang shares full bot prompt that auto-blocks travel time on your calendar — alexandr_wang · 2026-09-19
- Dev declares 'Long live AGENTS.md' as the agent spec keeps gaining traction — _jaydeepkarale · 2026-09-19
- Claude Code user shares tip: ask the agent to reorganize your session sidebar — steipete · 2026-09-19