Muse Glimmer Hits 230 tok/s on a Single RTX 5090 via SGLang
BanghuaZ · x · 2026-08-11
LMSYS shared a demonstration of Meta's newly open-sourced Muse Glimmer (30B dense model) achieving impressive local inference speeds.
Powered by the SGLang framework with NVFP4 and DFlash optimizations, the model reaches around 230 tokens/s on a single RTX 5090. It also works out of the box on RTX Pro 6000, DGX Spark, and Mac (MLX). This combination of speed and reliability makes building blazing-fast, always-on local agents highly feasible.
More from Infra
- Inside Xanadu's Lab: Ultra-low Loss Thin Film Lithium Niobate Switch Wafers — ceciletamura · 2026-08-11
- Beyond GPUs: Rethinking the Energy and Architecture Stack for Next-Gen AI Inference — prateekj · 2026-08-11
- Running Local LLMs on Strix Halo: Are 64GB/128GB RAM Variants Practical? — riklaunim · 2026-08-11
- Open Models Matching Cloud? It's Now an Engineering Tradeoff — cocktailpeanut · 2026-08-11
- GitHub Actions Outage Last Week: Users Await Incident Report — SkyLi0n · 2026-08-11
- Wall Street Giants Partner with Nvidia on $500B AI Infrastructure Financing — firstadopter · 2026-08-11