MOSS-VL Quantized Models Run Locally on 24GB VRAM for Multimodal Tasks

huggingface · x · 2026-08-12

OpenMOSS released FP8 and NF4 quantized versions for the MOSS-VL model series, supporting image, video, and real-time streaming understanding. MOSS-VL-Instruct is optimized for local inference and batch processing, while MOSS-VL-Realtime targets continuous video analysis for cameras and livestreams. The quantized models can run stably within 24GB VRAM, significantly lowering the hardware barrier for local multimodal deployment.

Original post →

More from Infra

Infra channel →