MOSS-VL: Open Vision-Language Model Family Enabling Real-Time Interaction

OpenMOSS-Team · hf · 2026-08-18

The OpenMOSS team released the technical report for MOSS-VL, an open vision-language model family. It enables real-time interaction by attending to visual information via gated cross-attention during generation. Utilizing a synthesized interaction corpus and staged curriculum, it achieves strong streaming performance with reduced time-to-first-token latency.

Original post →

More from Multimodal

Multimodal channel →