RefCaptioner: Precise Multi-Reference Image Grounding in Video Captioning
机器之心 · wechat · 2026-08-11
Peking University and the Kling team jointly proposed RefCaptioner, addressing the sharp decline in VLMs' ability to map video objects to multiple reference images.
- The Challenge: Existing video captioning models struggle with multiple reference images, often blindly citing irrelevant images or failing to group multi-view photos of the same subject.
- RefCaptioner Model: Built on Qwen3-VL-8B, it uses mixed-data SFT to preserve general captioning skills. It introduces HCD-GRPO, a dual-branch reward function featuring factuality rewards, Distractor-Aware Evidence Suppression (DAES), and Cross-Reference Semantic Coherence (CRSC) to prevent the model from cheating by simply spamming image tags.
- MRVBench Benchmark: The team open-sourced a parallel dataset of 462 videos with up to 22 candidate references, explicitly separating caption factuality from multi-reference grounding.
- Results: The 8B RefCaptioner achieves state-of-the-art performance among open-source models, approaching Gemini-3.1-Pro and surpassing GPT-5.4. Downstream video reconstruction tests confirm the high fidelity of its grounded captions.
More from Multimodal
- Grok Imagine Stuns Users by Converting Generated Images to Video In-Platform — XFreeze · 2026-08-11
- First Fully AI-Generated Feature Film Open-Sources All 473,600 Assets for $2M — LinusEkenstam · 2026-08-11
- AI Video Fail: Recreating the Iconic Miami Vice Scene with Bert and Ernie — dreamwieber · 2026-08-11
- 10-Minute AI-Generated Film Goes Viral with 300K Views in a Day — ZabihullahAtal · 2026-08-11
- One Stylus Tap Turns Raw Ingredients Into Gourmet Ramen Using AI Video — SimplyAnnisa · 2026-08-11
- How to Identify AI Video Models by Their Default Duration — cocktailpeanut · 2026-08-11