Ollama v0.40.0 auto-runs supported models on Apple's MLX runtime

lmoroney · x · 2026-10-07

Ollama v0.40.0 (released Tuesday) automatically runs any model architecture supported by the MLX runtime on Apple Silicon. MLX is Apple's open-source ML framework built around unified memory in M-series chips. Release notes list Qwen 3.8, Gemma 4, Qwen 3.6, Qwen 3.5, the Nimble, Tev1, Clef and Clef Flash decision models, and EmbeddingGemma 2, with more being tested. No speed numbers were published — the author suggests measuring yourself by comparing eval rates via ollama run --verbose before and after updating.

Original post →

More from Infra

Infra channel →