Meta Open-Sources Muse Glimmer: 30B Agentic Model That Fits a 24GB Consumer GPU
bibryam · x · 2026-09-26
Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter model open-sourced under Apache 2.0, designed for always-on local agent workflows.
- Deployment: its 17GB quantization fits a single 24GB consumer GPU with 1% average degradation across 15 benchmarks; optimized llama.cpp, MLX and ExecuTorch integrations are landing soon.
- Focus: local agents, function calling, local coding and LLM-as-a-judge evaluation — fully offline, no cloud required.
- Weights are on Hugging Face with developer documentation; the blog details how capabilities like long-horizon task execution were trained together.
It's Meta's latest bet that well-trained small models can approach frontier performance on targeted agentic tasks.
More from Infra
- Running Qwen3.8 Locally on 6x3090s with exllamav3 Hits 80-120 tok/s — takoulseum · 2026-09-26
- Gemma 4 31B UD-Q8_K_XL Suddenly Hits "Exceeds Shared Memory" at 68k Context — Few_Professional6859 · 2026-09-26
- What happens when data centers must replace hundreds of billions in GPUs? — No-Papaya-9289 · 2026-09-26
- SGLang team launches $ARKIE memecoin that routes 2% trade fees into GPUs for open AI infra — BanghuaZ · 2026-09-26
- Google confirms Gemini 4 is in early post-training as it readies to send TPUs to space — 量子位 · 2026-09-26
- TSMC approves $60.7bn in capital appropriations in 91 days, betting big on advanced nodes — julsimon · 2026-09-26