Muse Glimmer Hits 230 tok/s on a Single RTX 5090 via SGLang

BanghuaZ · x · 2026-08-11

LMSYS shared a demonstration of Meta's newly open-sourced Muse Glimmer (30B dense model) achieving impressive local inference speeds.

Powered by the SGLang framework with NVFP4 and DFlash optimizations, the model reaches around 230 tokens/s on a single RTX 5090. It also works out of the box on RTX Pro 6000, DGX Spark, and Mac (MLX). This combination of speed and reliability makes building blazing-fast, always-on local agents highly feasible.

Related event: Extreme Local Inference: Single GPUs Run 30B Models with Massive Context and High TPS(8 posts)→

Original post →

More from Infra

Infra channel →