VideoChat3 debuts as a fully open 4B video model for long and streaming clips

机器之心 · wechat · 2026-07-22

VideoChat3 brings fully open video understanding to 4B scale

Researchers from Nanjing University, Shanghai AI Lab, NTU, and Peking University released VideoChat3, a 4B-parameter multimodal model for general video understanding.

Main ideas

Data and training

The team also open-sourced three datasets:

Training is done in four stages, from vision pretraining to long-video and streaming instruction tuning, using 25M samples in total.

Reported results

VideoChat3 says it matches or beats other open models across long-video and streaming benchmarks, including strong gains over Qwen3-VL-4B in direct comparisons. It also cuts latency and compute for longer inputs, with large wins at 2048 frames.

The project is fully open: data, code, and weights are all released.

Original post →

More from Multimodal

Multimodal channel →