开源引擎 Tensorfold 让 Mac mini M5 Pro 跑 Qwen3.8-27B 达 40-60 tps

AdRepulsive7837 · reddit · 2026-09-30

一位用户实测开源推理引擎 Tensorfold:在 Mac mini M5 Pro(64GB)上用官方 4bit 量化模型 Qwen3.8-27B-MLX-4bit 配合 z-lab/Qwen3.8-27B-DFlash2 草稿模型,达到 40-60 tok/s,称这是 Mac 生态里首个速度超过 MLX 系(MTPLX)的推理引擎——他此前测试过 omlx、dflash2、mlx、llama-cpp、lm-studio、unsloth 均不敌 MTPLX。作为对比,其 RTX-3090ti 跑 GSQ-RCO-GGUF(MTP IQ3S,12.1GB)为 50-70 tok/s,水平接近,但 3090ti 的 prefill 快约 4 倍。作者因此认真考虑用 Mac mini 替代 3090ti 作为主力推理服务器。

原文链接 →

「Infra」频道最新

更多「Infra」频道 AI 资讯 →