RefCaptioner: Grounding Video Captions to Multiple Reference Images

KlingTeam · hf · 2026-07-31

Existing video captioning models cannot explicitly ground local visual elements to multiple reference images. To address this, researchers introduced the task of multi-reference image-grounded video captioning and proposed RefCaptioner.

Original post →

More from Multimodal

Multimodal channel →