Running Meta's Muse Glimmer 30B on MacBook: Ollama MLX Hits 29 tok/s

ollama · x · 2026-08-11

A developer tested running Meta's newly released Muse Glimmer 30B model locally on a MacBook (M3 Max, 96GB), comparing available serving options.

Results show that the fastest method currently is Ollama's MLX engine (DFlash included), reaching 29 tokens/sec. Tuned llama.cpp runs at 21 tokens/sec, while raw mlx-vlm is not yet optimized and runs at 10 tokens/sec.

Original post →

More from Infra

Infra channel →