RefCaptioner: Grounding Video Captions to Multiple Reference Images
KlingTeam · hf · 2026-07-31
Existing video captioning models cannot explicitly ground local visual elements to multiple reference images. To address this, researchers introduced the task of multi-reference image-grounded video captioning and proposed RefCaptioner.
- Framework: A two-stage post-training framework combining mixed-data SFT with Hierarchical Coverage-Discounted GRPO to improve reference selection, phrase-level binding, and cross-reference consistency.
- Data & Benchmark: Constructed a corpus of 20,000 videos and 171,354 reference images, alongside MRVBench for evaluating caption factuality and multi-reference grounding.
- Performance: Achieves the best overall performance among open-source models while remaining competitive on standard video captioning benchmarks, highly preferred by human annotators.
More from Multimodal
- Hailuo AI Demonstrates 15-Second Video Generation from a Single Image — LudovicCreator · 2026-07-31
- MiniMax H3 Video Model Enters Chatbot Arena, Open Weights Coming Soon — arena · 2026-07-31
- Mind-Blowing AI Concept: Visualizing Pigeon 'Flight-Mile' Delivery to Dodge Traffic — sivakumaranvlsi · 2026-07-31
- MiniMax H3 Video Model Now Available on Vercel AI Gateway — evilrabbit_ · 2026-07-31
- Testing Flux 3: Generating Synchronized Split-Screen Videos via Complex Prompts — umesh_ai · 2026-07-31
- MPIE-Bench: Evaluating Anatomical Errors in Multi-Person Image Editing — muset-ai · 2026-07-31