RefCaptioner: Precise Multi-Reference Image Grounding in Video Captioning

机器之心 · wechat · 2026-08-11

Peking University and the Kling team jointly proposed RefCaptioner, addressing the sharp decline in VLMs' ability to map video objects to multiple reference images.

Original post →

More from Multimodal

Multimodal channel →