VideoChat3: 4B Parameter Open Video MLLM Outperforms Larger Models
jiqizhixin · x · 2026-07-31
Researchers from Nanjing University, Shanghai AI Lab, and other institutes introduced VideoChat3, a fully open-source video multimodal large language model (MLLM) designed to efficiently process everything from short clips to hour-long streams.
Core Technical Highlights:
- Utilizes a 3D Vision Transformer and adaptive frame processing to understand videos efficiently without melting the GPU.
- Trained on three massive curated datasets covering general, long-form, and streaming content.
Performance & Availability:
- With only 4 billion parameters, it outperforms prior open-source models of equal or larger sizes across general, long-form, and streaming benchmarks.
- Fully open-source: model weights, training code, and datasets are all publicly available.
More from Multimodal
- Creators Generate Entire 'Alligator' Music Video Using Seedance 2.0 — AIandDesign · 2026-07-31
- Higgsfield Teases Seedance 2.5: Setting a New Standard for AI-Generated UGC — hey_abusiddik · 2026-07-31
- Suno Introduces AI Cover Art Generation with Multilingual Support — suno · 2026-07-31
- Mediabunny v1.52.2 Adds Bicubic Interpolation and Mipmaps for Better Video Resizing — Vjeux · 2026-07-31
- Help Needed: Lossless Outpainting Workflow in ComfyUI for Ultrawide Rescale — Tasty_Welcome_665 · 2026-07-31
- AI-Generated 3D Animation Shows Highly Coordinated Expressions — xiaohu · 2026-07-31