Red Hat AI Releases Muse-Glimmer 30B FP8 Quantized Checkpoint, Halving Memory
vllm_project · x · 2026-08-10
Red Hat AI has released a quantized checkpoint for Meta's multimodal model Muse-Glimmer 30B, named Muse-Glimmer-30B-FP8-block. By applying block-wise FP8 quantization to the linear layers while keeping the vision tower in full precision, this version roughly halves the memory and disk footprint without losing significant capabilities. The model was quantized using LLM Compressor and is ready to be served via the vLLM project.
More from Infra
- Muse Glimmer Hits 230 tok/s on a Single RTX 5090 via SGLang — BanghuaZ · 2026-08-11
- SGLang v0.5.17 Released: Adds Support for Kimi K3 and MiniMax-H3 Video Generation — BanghuaZ · 2026-08-11
- Karpathy: Hybrid Setup of Cloud Executive Intelligence and Local Models is Very Appealing — karpathy · 2026-08-11
- Local Video Gen with MiniMaxH3: Workflow and Hardware Upgrade Notes — Last-Pie8057 · 2026-08-10
- Running Large Models on 8GB VRAM? Understanding Shared Memory in Local Deployment — Leary_2844 · 2026-08-10
- SemiAnalysis Deep Dive: Can TileRT Software Make NVIDIA GPUs Compete with Cerebras and Groq? — dylan522p · 2026-08-10