MLX-Serve v26.9.5 lands with Qwen-Image 2.1 and 4-way MTP streams at up to 122 tok/s on M4 Max
TheMoonMidas · x · 2026-09-23
- MLX-Serve v26.9.5 (Bonsai) is out, now with a companion iOS App, enabling local model serving on Apple Silicon.
- Image generation: Qwen-Image-2.1 text-to-image, image-to-image and negative-prompt support via 8bit (32GB Macs) and 4bit (16GB Macs) quantized packs; low-RAM machines load the text encoder per request and free it before denoising.
- Prism Bonsai 2: a Hadamard-rotated 2-bit Qwen3.8-27B (text + vision), billed as probably the fastest implementation, loads like any Qwen 27B.
- Concurrency: four MTP streams share one verify forward, lifting aggregate throughput from 72 to 119 tok/s on M4 Max; KV-copy fix boosts 4 streams at 28k context from 26 to 64 tok/s; with a DFlash drafter, concurrent requests batch via the MTP head (64 to 122 tok/s).
- 8-bit KV matches bf16 KV speed at long context. Open source on GitHub (1.5k stars).
More from Infra
- Gemini's sandbox escape, California's kill switch order, and Apple's server return: one theme, control — Dapper-Tale-4021 · 2026-09-23
- China reportedly hits near-GPT-5.6 intelligence at 1/15th cost with $2.6M RL run — djcows · 2026-09-23
- Al Gore: all AI data centers emit less than the world's uncovered landfills — jeffclune · 2026-09-23
- Developer building infrastructure to run 1 million concurrent AI agents on Kubernetes — gonzohst1 · 2026-09-23
- LTX-2.5 on a rented 96GB Blackwell: 10s of 1080p video with audio in 231s for 7 cents — Realistic-Fennel-190 · 2026-09-23
- Supermicro Now Shipping NVIDIA Vera Rubin NVL72 Racks — AIFlow_ML · 2026-09-23