Tensorfold hits 40-60 tok/s on Qwen3.8-27B with Mac mini M5 Pro, beating MLX

AdRepulsive7837 · reddit · 2026-09-30

A user reports the open-source inference engine Tensorfold reaches 40-60 tok/s running Qwen3.8-27B (4-bit, with a DFlash2 drafting model) on a Mac mini M5 Pro 64GB — the first Mac engine to beat MLX in their tests of omlx, dflash2, mlx, llama-cpp, lm-studio and unsloth. Their RTX-3090ti manages 50-70 tok/s with 4x faster prefill, and they now consider replacing the GPU entirely.

Original post →

More from Infra

Infra channel →