Qwen on M2 Ultra: latest oMLX update brings substantial local inference speedup
Thrumpwart · reddit · 2026-09-13
A Reddit user shares an update on running Qwen 3.8 Flash locally on an M2 Ultra: the latest oMLX release introduces a substantial speedup, with a benchmark screenshot attached. A notable signal for Mac-based local inference performance.
More from Infra
- A silent CPU shortage is hitting cloud reservations for mid-size companies — Sethwinterroth · 2026-09-13
- DeepSeek V4.1 Flash Hits 40 tok/s on 8x A40 with Open-Source TensorSharp Engine — fuzhongkai · 2026-09-13
- Samsung's first 2nm Taylor fab fully booked by Tesla, Arm and Broadcom pre-launch — Beth_Kindig · 2026-09-13
- Reddit theory: the AI 'slowdown' talk masks a coming compute, power and datacenter crunch — maxpayne07 · 2026-09-13
- Team With 2x A100 Grant Open-Sources Roadmap for Distilling Offline Edge Models — Lerok-Persea · 2026-09-13
- Prediction: hyperscalers will tighten AI capex in 2027 — pdamodaran · 2026-09-13