Fully Open-Source Video MLLM VideoChat3

MCG-NJU · hf · 2026-07-17

VideoChat3 is introduced as a fully open, more efficient, and highly generalizable video MLLM.

Main Goals

Common issues with existing open-source video models:

Method Design

The authors improved upon these from two directions:

1) Enhancing Efficiency

2) Enhancing Performance

Built a scalable video data synthesis pipeline, organizing three training sets:

Covering general, long video, and streaming video scenarios.

Results

On general, long video, and streaming video benchmarks, VideoChat3 (with 4B parameters) achieved better results and higher efficiency than previous open-source models of the same or even larger scales.

Related event: Fully Open-Source Video MLLM VideoChat3 Released(3 posts)→

Original post →

More from Multimodal

Multimodal channel →