MLX-Serve 26.10.1 ships with up to 66% faster Qwen3.8 27B inference on Apple Silicon
TheMoonMidas · x · 2026-10-01
MLX-Serve 26.10.1 delivers broad speedups with byte-identical outputs across 18 models.
- Speculative decoding: Qwen3.8 27B with its drafter gains +66% on M5 Ultra, +28% on M4 Max, +37% on M1 Pro
- Qwen Flash Next: +11% decode on M4 Max, +51% prefill on M5 Ultra (up to 5400 PP)
- Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster
- New: GGUF models now run on a custom MLX engine written in Zig and Metal (experimental, replacing the llama.cpp fallback); direct support for Sushi-format Flash Next packs; Qwen-Image 2.1 instruction-based photo editing; drafting enabled by default for all supported models
More from Infra
- A goldmine resource for learning GPU programming internals — goyal__pramod · 2026-10-02
- Gradium claims fastest TTS yet with ~50ms time-to-first-audio, tops sub-100ms naturalness — mattturck · 2026-10-02
- Engineer: Gemini 4 Argon drives fleet-wide data center optimization gains — rakyll · 2026-10-02
- Aleph Alpha Details MoE Pre-training Scaling: 30B-A3B on 16-512 B200s at 35% MFU — CatAstro_Piyush · 2026-10-02
- HBM to eat 22% of world's DRAM wafers as memory prices triple in six months — aronchick · 2026-10-02
- Can Google's AI chip beat Nvidia? A video analysis — Ill_Vegetable169 · 2026-10-02