Running Meta's Muse Glimmer 30B Locally on Mac Studio at 30 tok/s
MaziyarPanahi · x · 2026-08-10
A developer has successfully run Meta's newly released Muse Glimmer 30B multimodal model (GGUF format) locally on a Mac Studio.
A recorded demo shows a real two-turn conversation where the model streams its reasoning live before landing on the answer, measuring about 30 tokens/second via the local API. The author is currently testing the local stack across healthcare workflows, visual pipelines, tool use, and full agentic tasks.
Related event: Meta's Muse Glimmer 30B Runs Locally on Mac Studio with Day-Zero Support(2 posts)→
More from Models
- Running 30B Model on Single RTX 5060 Ti with 131k Context Window — BazzyIm · 2026-08-10
- Topology Test: Grok Imagine 2.0 Beats ChatGPT by Correctly Understanding Genus — luismbat · 2026-08-10
- Testing LLM Reasoning: A Solar Eclipse Probability Math Problem — Afinetheorem · 2026-08-10
- OpenAI Locks Down Astra After Model Raises First-Ever Critical Cyber Capability Fears — sksarkpoes3 · 2026-08-10
- Opinion: Does Google Need a Top-Tier Coding Model? Skipping to Gemini 4 — haider1 · 2026-08-10
- Kimi K3 Hits 295 tok/s in Tests, Confirmed Full Precision Without Quantization — JiaZhihao · 2026-08-10