MOSS-VL-Realtime is Now Open Source
victormustar · x · 2026-07-14
MosiAI has open-sourced **MOSS-VL-Realtime**, an 11B vision-language model designed for real-time visual understanding of continuous video streams under the Apache-2.0 license. It supports asking questions at any point during the video and features a "watch and answer" capability: it can correct or interrupt its answers when the scene changes, and remain silent when evidence is insufficient. The project also offers Base, Instruct, and Realtime versions, supporting bilingual (Chinese and English) multimodal understanding with a context window length of 256K tokens.
Related event: MOSS-VL-Realtime Open-Sourced for Real-time Vision(2 posts)→
More from Multimodal
- Anatomy of Dynamic AI Images: Subject, Environment, and Camera — GPU_FieldNotes · 2026-07-21
- MiniCPM-V 4.6 now runs locally on iPhone with no cloud dependency — amos_gyamfi · 2026-07-21
- Why AI action images still look static unless pose, motion and camera angle all work together — Jaded-Term-8614 · 2026-07-21
- Creator says they no longer shoot with a camera, but with prompts — taherdhanera · 2026-07-21
- PixVerse demo turns into a full sci-fi dark comedy set on Mars — aliscodes · 2026-07-21
- Alibaba’s Qwen-Audio-3.0-TTS-Plus takes #1 on Artificial Analysis Speech Arena — airesearch12 · 2026-07-21