Running Meta's Muse Glimmer 30B on MacBook: Ollama MLX Hits 29 tok/s
ollama · x · 2026-08-11
A developer tested running Meta's newly released Muse Glimmer 30B model locally on a MacBook (M3 Max, 96GB), comparing available serving options.
Results show that the fastest method currently is Ollama's MLX engine (DFlash included), reaching 29 tokens/sec. Tuned llama.cpp runs at 21 tokens/sec, while raw mlx-vlm is not yet optimized and runs at 10 tokens/sec.
More from Infra
- RTX 5090 vs. Dual 48GB GPUs: A Hardware Upgrade Guide for Local AI Video Generation — Ammoryyy · 2026-08-11
- Self-Hosted Coding Agent in MicroVM Sandboxes with Local Inference and iOS App — tom_doerr · 2026-08-11
- SanDisk CEO Says Mid-80s Gross Margin Is a Fair Return for Storage Products — Beth_Kindig · 2026-08-11
- Jensen Huang says Nvidia is making AI compute an investable asset class — Polymarket · 2026-08-11
- Apple Approves tinygrad eGPU Driver: Macs Can Finally Use AMD/NVIDIA GPUs — ns123abc · 2026-08-11
- Tech Giants' Compute Budgets Eclipse US Federal Spending—We Live in Cyberpunk Now — tszzl · 2026-08-11