PANORAMA Grounds Every Caption Phrase to Pixel-Level Masks, Tops New PanoCaps Benchmark

Panorama-grounding · hf · 2026-09-17

Current VLMs produce fluent captions but struggle to reliably tie them to pixels. PANORAMA tackles panoptic grounded captioning: describing foreground objects and background regions while grounding every referring phrase with pixel-level masks.

PANORAMA achieves the best overall grounding on PanoCaps and matches or beats specialized models across pixel-level grounding tasks. Code, data and models are released.

Original post →

More from Multimodal

Multimodal channel →