Open MOSS Open-Sources Real-Time Multimodal Model
AdinaYakup · x · 2026-07-19
Open MOSS has open-sourced MOSS-VL-Realtime, which is now available on Hugging Face. This 11B model family supports text, single/multiple images, single/multiple videos, and interleaved image-text inputs, covering both Chinese and English scenarios.
The project highlights several design choices: using Cross-Attention to separate visual encoding from language reasoning, applying XRoPE for unified spatiotemporal positional encoding, providing a unified conversation template applicable for offline/streaming/real-time interactions, and supporting a 256K token context. It can also continue processing new frames while generating a response, modifying or interrupting its answer based on scene changes, and even remaining silent when evidence is insufficient.
More from Multimodal
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- Dev builds interactive 3D product experience with GPT-6 Astra + Hyper3D Rodin — nikola_mr64990 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11