Fully Open-Source Video MLLM VideoChat3
MCG-NJU · hf · 2026-07-17
VideoChat3 is introduced as a fully open, more efficient, and highly generalizable video MLLM.
Main Goals
Common issues with existing open-source video models:
- Limited generalization, working well only in a few scenarios
- High computational overhead, making scaling difficult
- Partial openness, with incomplete training code/strategies/datasets affecting reproducibility
Method Design
The authors improved upon these from two directions:
1) Enhancing Efficiency
- Introduced Inflated 3D Vision Transformer (I3D-ViT)
- Adopted Adaptive Frame Resolution for streaming video perception
- Reduced the cost of processing video inputs during training and inference
2) Enhancing Performance
Built a scalable video data synthesis pipeline, organizing three training sets:
- VideoChat3-Academic2M
- VideoChat3-LV116K
- VideoChat3-OL617K
Covering general, long video, and streaming video scenarios.
Results
On general, long video, and streaming video benchmarks, VideoChat3 (with 4B parameters) achieved better results and higher efficiency than previous open-source models of the same or even larger scales.
Related event: Fully Open-Source Video MLLM VideoChat3 Released(3 posts)→
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11