Meta Releases Muse Glimmer: A 30B Open-Weight Local Agent Model

hsu_byron · x · 2026-08-10

Meta has officially released Muse Glimmer, a 30B-parameter dense model optimized for local, always-on agent workflows. Released under a permissive Apache 2.0 license with open weights, the model is designed to run entirely on consumer hardware like Macs or PCs with performant GPUs, delivering strong performance on key agentic use cases.

LMSYS subsequently shared benchmark data: with day-0 SGLang support, Muse Glimmer achieves around 230 tok/s on a single RTX 5090 using NVFP4 + DFlash. It also works out of the box on NVIDIA RTX Pro 6000, DGX Spark, and MLX for Mac. This blazing-fast and reliable local inference capability removes the speed blockers that previously hindered local agent development.

Related event: Meta Releases Muse Glimmer 30B Open-Source Model for Local Agents(30 posts)→

Original post →

More from Infra

Infra channel →