Meta Releases Muse Glimmer: A 30B Open-Weight Local Agent Model
hsu_byron · x · 2026-08-10
Meta has officially released Muse Glimmer, a 30B-parameter dense model optimized for local, always-on agent workflows. Released under a permissive Apache 2.0 license with open weights, the model is designed to run entirely on consumer hardware like Macs or PCs with performant GPUs, delivering strong performance on key agentic use cases.
LMSYS subsequently shared benchmark data: with day-0 SGLang support, Muse Glimmer achieves around 230 tok/s on a single RTX 5090 using NVFP4 + DFlash. It also works out of the box on NVIDIA RTX Pro 6000, DGX Spark, and MLX for Mac. This blazing-fast and reliable local inference capability removes the speed blockers that previously hindered local agent development.
Related event: Meta Releases Muse Glimmer 30B Open-Source Model for Local Agents(30 posts)→
More from Infra
- AMD Acquires Startup to Burn AI Weights Directly Into Silicon Chips — lemire · 2026-08-10
- Muse Spark 1.2 Sparse Model Optimized for Enterprise Hardware to be Open-Sourced — jack_w_rae · 2026-08-10
- llama.cpp Adds Day 0 Support for Muse Glimmer Model — jacek2023 · 2026-08-10
- Intel Announces $15B Stock Offering to Capitalize on AI Compute Demand — ryanshrout · 2026-08-10
- Hybrid Cloud + Local Model Architecture Shows Promise; Muse Spark 1.2 Open-Source Release Imminent — jack_w_rae · 2026-08-10
- 30B Model Hits 114 tps with Tuned Quants, Targeting <16G VRAM — tokenbender · 2026-08-10