Meta Returns to Open Source: Muse Glimmer Hits 230 tok/s on RTX 5090 via SGLang

ying11231 · x · 2026-08-11

Meta has released Muse Glimmer (a 30B dense open-weights model), featuring day-0 support from SGLang.

Utilizing SGLang and quantization, the model achieves around 230 tokens/s on a single RTX 5090. Quantized checkpoints for both Apple Silicon and NVIDIA hardware are now available on Hugging Face. This removes the speed and reliability blockers for local agents, running out-of-the-box on RTX Pro 6000, DGX Spark, and Mac (MLX).

Related event: Meta Open-Sources 30B Agentic Model Muse Glimmer(63 posts)→

Original post →

More from Infra

Infra channel →