开源引擎 Inco Splash 让 Qwen3.8-27B 在 M5 Max 跑出 144 tok/s

ResearchCrafty1804 · reddit · 2026-09-19

开源推理引擎 Inco Splash 针对 Apple Silicon 定制,宣称 Qwen3.8-27B 在 M5 Max MacBook Pro 上达 144 tok/s:解码速度约为 Ollama 的 3 倍、oMLX 的 2 倍,agent 扇出子 agent 时接近 4 倍。要求 M3 及以上芯片、macOS 26.4+、36GB 内存。一行 brew install + splash serve 即可部署,兼容 Claude Code、OpenCode、Codex 等 agent,也可在 LM Studio 中作为 runtime 使用。

所属事件:开源引擎 Inco Splash 助 Apple Silicon 本地跑大模型提速(3 条相关)→

原文链接 →

「编程与Agent」频道最新

更多「编程与Agent」频道 AI 资讯 →