Open MOSS Open-Sources Real-Time Multimodal Model
AdinaYakup · x · 2026-07-19
Open MOSS has open-sourced **MOSS-VL-Realtime**, which is now available on Hugging Face. This 11B model family supports text, single/multiple images, single/multiple videos, and interleaved image-text inputs, covering both Chinese and English scenarios. The project highlights several design choices: using **Cross-Attention** to separate visual encoding from language reasoning, applying **XRoPE** for unified spatiotemporal positional encoding, providing a unified conversation template applicable for offline/streaming/real-time interactions, and supporting a **256K token** context. It can also continue processing new frames while generating a response, modifying or interrupting its answer based on scene changes, and even remaining silent when evidence is insufficient.
More from Multimodal
- Reddit users say Krea 2 Turbo regains strong facial expressions with bypass LoRAs — YentaMagenta · 2026-07-21
- AI music demo blends Suno v5.5, Reason Studios and Grok Imagine 1.5 — Kyrannio · 2026-07-21
- AI anime workflow article breaks storytelling into repeatable prompt steps — Aiden_Tech_Ai · 2026-07-21
- Claude is being pitched as a free workflow for viral YouTube Shorts scripts — Aiden_Tech_Ai · 2026-07-21
- A Bittensor game demo claims a two-person team built a playable 3D world in 30 days — markjeffrey · 2026-07-21
- Gemini Omni Flash turns a boat cabin into a cave inside Flow — chrisfirst · 2026-07-21