Ollama 0.40 RC Tested: MLX Speed Is Real, But the Stable Version Already Has It
TheOyinbooke · x · 2026-09-29
A hands-on comparison of Ollama 0.40.0-rc0 (MLX-by-default on Apple Silicon) vs stable 0.34.4 on a 16GB Mac Mini M4: MLX delivers up to 2.5x faster Gemma 4 runs, but the stable version achieves the same speed via a different model tag. The pre-release also swapped in a 27.8B nvfp4 MLX build of Qwen 3.8 that downloads 18GB then fails to load on 16GB (needs 16.9 GiB vs 11.3 GiB available). Verdict: 5/10 — right idea, changes almost nothing today. Tested with 4 fixed prompts, 3 runs each, median reported.
More from Infra
- Dagger founder: the Great CI Bottleneck of 2026 is a software problem, not hardware — msharmas · 2026-09-29
- LayerSkip: Self-Speculative Decoding Speeds Up LLMs Without a Draft Model — burkov · 2026-09-29
- SpaceX outlines supercomputer, Terafab, Gigasat and new Louisiana Starbase plans — elonmusk · 2026-09-29
- Musk says space will hold nearly all compute; Google tests if TPUs work there — CackleRooster · 2026-09-29
- 95+ TPS and 262K context for Qwen 27B on a single RTX 3090 with LlamAmpere v0.4 — Brief-Tap-6616 · 2026-09-29
- More Americans oppose a local data center than a nuclear reactor, says Cathie Wood — PeterDiamandis · 2026-09-29