vLLM Releases Deployment Guide for Muse-Glimmer-30B Local Multimodal Agent

vllm_project · x · 2026-08-10

vLLM Recipes has published a deployment guide for the Muse-Glimmer-30B model. This is a dense 29.6B parameter vision-language model featuring a ViT-G/14 perception encoder and 128K context, designed for local agentic tasks on consumer hardware.

A unique aspect of the model is its output format: it avoids JSON tool calls and <think> tags for reasoning. Instead, it uses channel-scoped message sequences and XML-style ATEM tool calls. Deployment requires the dedicated museglimmer tool-call and reasoning parsers. Multiple weight variants are available, including BF16, NVFP4 quantization, and a draft head for speculative decoding.

Related event: Meta Releases Muse Glimmer 30B Open-Source Model for Local Agents(30 posts)→

Original post →

More from Infra

Infra channel →