Meta Returns to Open Source: Muse Glimmer Hits 230 tok/s on RTX 5090 via SGLang
ying11231 · x · 2026-08-11
Meta has released Muse Glimmer (a 30B dense open-weights model), featuring day-0 support from SGLang.
Utilizing SGLang and quantization, the model achieves around 230 tokens/s on a single RTX 5090. Quantized checkpoints for both Apple Silicon and NVIDIA hardware are now available on Hugging Face. This removes the speed and reliability blockers for local agents, running out-of-the-box on RTX Pro 6000, DGX Spark, and Mac (MLX).
Related event: Meta Open-Sources 30B Agentic Model Muse Glimmer(63 posts)→
More from Infra
- Micron Exec: AI Customer Roadmaps Now Visible Beyond 2030 — BenBajarin · 2026-08-11
- Vercel Makes Sandbox Egress Firewall Free, Citing AI Agent Network Escape Risks — cramforce · 2026-08-11
- Mega Funds Form $500B Alliance to Keep Arm's Length from NVIDIA — annbordetsky · 2026-08-11
- RTX 5090 vs. Dual 48GB GPUs: A Hardware Upgrade Guide for Local AI Video Generation — Ammoryyy · 2026-08-11
- Self-Hosted Coding Agent in MicroVM Sandboxes with Local Inference and iOS App — tom_doerr · 2026-08-11
- SanDisk CEO Says Mid-80s Gross Margin Is a Fair Return for Storage Products — Beth_Kindig · 2026-08-11