Open MOSS Open-Sources Real-Time Multimodal Model
AdinaYakup · x · 2026-07-19
Open MOSS has open-sourced MOSS-VL-Realtime, which is now available on Hugging Face. This 11B model family supports text, single/multiple images, single/multiple videos, and interleaved image-text inputs, covering both Chinese and English scenarios.
The project highlights several design choices: using Cross-Attention to separate visual encoding from language reasoning, applying XRoPE for unified spatiotemporal positional encoding, providing a unified conversation template applicable for offline/streaming/real-time interactions, and supporting a 256K token context. It can also continue processing new frames while generating a response, modifying or interrupting its answer based on scene changes, and even remaining silent when evidence is insufficient.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11