Meta Releases Muse Glimmer: A 30B Open-Weight Agentic Model Running on 24GB VRAM
alexandr_wang · x · 2026-08-10
Meta Superintelligence Labs has introduced Muse Glimmer, a 30-billion-parameter open-weight agentic model released under the Apache 2.0 license.
- Local Execution: Optimized for always-on local agent workflows, it runs on a single consumer GPU with 24GB VRAM. It supports local coding, function calling, and LLM-as-a-judge evaluation.
- Inference Optimization: Built with a custom architecture, it quantizes weights to 4-bit to fit under 20GB. It uses a dflash drafter to propose token blocks verified in parallel by the main model, ensuring responsiveness.
- Agentic Capabilities: Functions as a fully capable agent with planning, tool calling, result checking, and failure recovery, rivaling much larger models.
- Ecosystem: Weights are available on Hugging Face, with integrations for Ollama, vLLM, llama.cpp, and others rolling out this week.
An open-weight version of another model, Muse Spark 1.2, is also coming soon.
Related event: Meta Releases Muse Glimmer 30B Open-Source Model for Local Agents(30 posts)→
More from Infra
- llama.cpp Adds Day 0 Support for Muse Glimmer Model — jacek2023 · 2026-08-10
- Intel Announces $15B Stock Offering to Capitalize on AI Compute Demand — ryanshrout · 2026-08-10
- Hybrid Cloud + Local Model Architecture Shows Promise; Muse Spark 1.2 Open-Source Release Imminent — jack_w_rae · 2026-08-10
- 30B Model Hits 114 tps with Tuned Quants, Targeting <16G VRAM — tokenbender · 2026-08-10
- DeepSeek API Costs $1.14/Day; Dual DGX Breaks Even in 24 Years — delduca · 2026-08-10
- Beyond Compute: AI Data Centers Face the Power Scarcity Bottleneck — ingliguori · 2026-08-10