oMLX 0.7.1.dev1 adds decision models and batched prefill, +31% Qwen3.6 decode speed on M5 Max
pcuenq · x · 2026-10-10
oMLX 0.7.1.dev1 brings notable updates for local inference on Apple Silicon:
- Speedups: Qwen3.6-35B-A3B decode on M5 Max hits 205 tok/s (+31%); GLM-5.3-Flash on M3 Ultra reaches 40.6 tok/s (+30%), with fused decode kernels now supporting M3/M4; Qwen3.8-Flash-Next with expert offload hits 636 tok/s (+61%) on 32K prefill.
- Batched prefill: concurrent prompts share one forward pass, cutting time-to-first-token from 7.41s to 5.75s (-22%) with 8 prompts.
- Decision models: new /v1/systemone endpoint serves Clef, Clef-Flash, and OpenJev, which return per-option probabilities for typed questions instead of generating text, with auto-detection and ready-made quantized checkpoints.
- Refreshed web UI; supports macOS 15+.
More from Infra
- Micron didn't end up with zero Rubin HBM, as analyst commentary backpedals — suchenzang · 2026-10-10
- System 1 ANE: millisecond AI decisions on Apple's Neural Engine with no text generation — pcuenq · 2026-10-10
- Bittensor GPU rental network sees 45% spend growth, 106% more rentals in monthly report — markjeffrey · 2026-10-10
- Microsoft doubles down on local AI with Nvidia RTX Spark Surface Ultra priced up to $5,899 — MooseEfficient2151 · 2026-10-10
- Building a pit crew for Grok Bot: frontier model plans, free models grind — alexcovo_eth · 2026-10-10
- AI boom turns into a debt boom: Oracle 5y CDS near record 261bps, implying 20.4% default odds — cyb3rops · 2026-10-10