SGLang Announces Day-0 Support for Meta's Muse Glimmer, Hitting 230 tok/s on RTX 5090
NVIDIAAI · x · 2026-08-10
SGLang has announced Day-0 support for Meta's newly released Muse Glimmer multimodal model.
With NVFP4 and DFlash optimizations enabled, the model achieves inference speeds of approximately 230 tokens/second on a single RTX 5090 GPU. It also runs out of the box on NVIDIA RTX PRO 6000, DGX Spark, and Apple Silicon via the MLX framework.
Related event: Meta Releases Open-Source Muse Glimmer 30B Model(100 posts)→
More from Infra
- Single used GPU matches Opus on internal workloads, local AI underestimated — AccBalanced · 2026-08-12
- NVIDIA Nemotron 3.5 Lightning Goes Live on CoreWeave Serverless — wandb · 2026-08-12
- AWS Releases Reference Architecture for Enterprise Claude Apps Gateway — AWS ML Blog · 2026-08-11
- NVIDIA Expert: Multi-Token Techniques Become Day-Zero Norm for Inference — PavloMolchanov · 2026-08-11
- Autonomous Computer: $26K Dual RTX 5090 Workstation Targets Local Frontier Models — dee_hw · 2026-08-11
- NVIDIA Nemotron 3.5 Lightning Goes Live on Crusoe for High-Volume Agent Inference — Scobleizer · 2026-08-11