SGLang Announces Day-0 Support for Meta's Muse Glimmer, Hitting 230 tok/s on RTX 5090

NVIDIAAI · x · 2026-08-10

SGLang has announced Day-0 support for Meta's newly released Muse Glimmer multimodal model.

With NVFP4 and DFlash optimizations enabled, the model achieves inference speeds of approximately 230 tokens/second on a single RTX 5090 GPU. It also runs out of the box on NVIDIA RTX PRO 6000, DGX Spark, and Apple Silicon via the MLX framework.

Related event: Meta Releases Open-Source Muse Glimmer 30B Model(100 posts)→

Original post →

More from Infra

Infra channel →