MLX-Serve 27B gets speculative branching, up to 32% faster generation on M4 Max
TheMoonMidas · x · 2026-09-27
Developer ddalcu shares that MLX-Serve now incorporates tricks from TensorFold: it makes several branching guesses for the next tokens and verifies them all in a single pass. On an M4 Max running a 27B model, this yields up to 20% faster generation than TensorFold and 32% versus version 26.9.5 — a notable local-inference throughput gain for the Apple MLX ecosystem.
More from Infra
- AI agents' PR flood pushed this team from GitHub runners to on-prem CI — peterjliu · 2026-09-28
- Homelab upgrade: 2x RTX 5090 with NVFP4 + speculative decoding beats 3x RTX 3090 — CompetitiveDraft9381 · 2026-09-28
- DeepSeek's DSec paper: 3M sandboxes/day, 5,000 creations per second on 160 nodes — petrusenko_max · 2026-09-28
- Is Switching from llama.cpp to vLLM Worth It? Day-0 Model Support Is the Draw — Exciting-Engine882 · 2026-09-28
- PyTorch Conference to cover enterprise-grade agentic inference with PyTorch and vLLM — PyTorch · 2026-09-28
- Fully Local Animation with MiniMax H3 on an RTX 5070Ti — maciusik · 2026-09-28