Ollama v0.40.0 auto-runs supported models on Apple's MLX runtime
lmoroney · x · 2026-10-07
Ollama v0.40.0 (released Tuesday) automatically runs any model architecture supported by the MLX runtime on Apple Silicon. MLX is Apple's open-source ML framework built around unified memory in M-series chips. Release notes list Qwen 3.8, Gemma 4, Qwen 3.6, Qwen 3.5, the Nimble, Tev1, Clef and Clef Flash decision models, and EmbeddingGemma 2, with more being tested. No speed numbers were published — the author suggests measuring yourself by comparing eval rates via ollama run --verbose before and after updating.
More from Infra
- AI chip testing is the new bottleneck: winstek approves NT$13.4B capex on NT$2.7B revenue — tengyanAI · 2026-10-07
- Qwen3-TTS 97ms Streaming Had No Repro Code — Community Repo Fills the Gap in vLLM — vllm_project · 2026-10-07
- Project Maya runs 321B GLM-5.3-Flash locally on two old V100s at up to 40 tok/s — lxfater · 2026-10-07
- Data center water and power fears overblown? Aluminum smelter uses a Boston-sized grid — csuwildcat · 2026-10-07
- Omarchy Linux ships official browser-based remote desktop plugin — juntao · 2026-10-07
- Datology AI open-sources Zephon, a deterministic on-the-fly data loader built for massive ablations — pratyushmaini · 2026-10-07