Tensorfold hits 40-60 tok/s on Qwen3.8-27B with Mac mini M5 Pro, beating MLX
AdRepulsive7837 · reddit · 2026-09-30
A user reports the open-source inference engine Tensorfold reaches 40-60 tok/s running Qwen3.8-27B (4-bit, with a DFlash2 drafting model) on a Mac mini M5 Pro 64GB — the first Mac engine to beat MLX in their tests of omlx, dflash2, mlx, llama-cpp, lm-studio and unsloth. Their RTX-3090ti manages 50-70 tok/s with 4x faster prefill, and they now consider replacing the GPU entirely.
More from Infra
- Vital Optics reportedly lands 6-inch InP substrate orders from a leading global customer — zephyr_z9 · 2026-09-30
- DGX Spark Handbook: how a 128GB low-bandwidth box now matches cloud inference speed — mervenoyann · 2026-09-30
- Cloudflare rebuilds Containers for agent sandboxes: 6x faster startups, 648ms median — michellechen · 2026-09-30
- Autonomous launches 4x RTX 5090 local AI workstation at $43,900 with 128GB+ VRAM — dee_hw · 2026-09-30
- Dev warns: running agents that know your whole life on someone else's computer — SuhailKakar · 2026-09-30
- AI bots are aggressively mining library websites, straining open-access resources — _akpiper · 2026-09-30