Fully Open-Source Video MLLM VideoChat3
MCG-NJU · hf · 2026-07-17
VideoChat3 is introduced as a fully open, more efficient, and highly generalizable video MLLM.
Main Goals
Common issues with existing open-source video models:
- Limited generalization, working well only in a few scenarios
- High computational overhead, making scaling difficult
- Partial openness, with incomplete training code/strategies/datasets affecting reproducibility
Method Design
The authors improved upon these from two directions:
1) Enhancing Efficiency
- Introduced Inflated 3D Vision Transformer (I3D-ViT)
- Adopted Adaptive Frame Resolution for streaming video perception
- Reduced the cost of processing video inputs during training and inference
2) Enhancing Performance
Built a scalable video data synthesis pipeline, organizing three training sets:
- VideoChat3-Academic2M
- VideoChat3-LV116K
- VideoChat3-OL617K
Covering general, long video, and streaming video scenarios.
Results
On general, long video, and streaming video benchmarks, VideoChat3 (with 4B parameters) achieved better results and higher efficiency than previous open-source models of the same or even larger scales.
Related event: Fully Open-Source Video MLLM VideoChat3 Released(3 posts)→
More from Multimodal
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22
- Solo founder turns complaints on screen into bug reports with a local MCP server — phdptsd · 2026-07-22
- Google demo says Gemma 4 can inspect car damage from video in under 6 seconds — soumitrashukla9 · 2026-07-22