Red Hat AI Releases Muse-Glimmer 30B FP8 Quantized Checkpoint, Halving Memory

vllm_project · x · 2026-08-10

Red Hat AI has released a quantized checkpoint for Meta's multimodal model Muse-Glimmer 30B, named Muse-Glimmer-30B-FP8-block. By applying block-wise FP8 quantization to the linear layers while keeping the vision tower in full precision, this version roughly halves the memory and disk footprint without losing significant capabilities. The model was quantized using LLM Compressor and is ready to be served via the vLLM project.

Original post →

More from Infra

Infra channel →