PANORAMA Grounds Every Caption Phrase to Pixel-Level Masks, Tops New PanoCaps Benchmark
Panorama-grounding · hf · 2026-09-17
Current VLMs produce fluent captions but struggle to reliably tie them to pixels. PANORAMA tackles panoptic grounded captioning: describing foreground objects and background regions while grounding every referring phrase with pixel-level masks.
- PanoCaps benchmark: human-annotated from panoptic segmentation datasets, with dense near-complete pixel coverage and entity-level image-text alignment, plus a phrase-mask matching protocol and generalized Panoptic Quality (gPQ) metric
- Method: formulates grounding as selecting from a phrase-conditioned pool of mask proposals, conditioning a pretrained segmenter on contextualized phrase representations, trained jointly with captioning
PANORAMA achieves the best overall grounding on PanoCaps and matches or beats specialized models across pixel-level grounding tasks. Code, data and models are released.
More from Multimodal
- Midjourney for style, Gemini Nano Banana for edits: a two-step AI art workflow — michaelrabone · 2026-09-17
- Midjourney --sref 3530132817 turns any character into cute chibi caricature art — michaelrabone · 2026-09-17
- AI-generated rooftop photoshoot for a ceramics brand stuns with product-grade visuals — sidahuj · 2026-09-17
- Why is there still no SOTA open-weights image editing model? — Rheumi · 2026-09-17
- 34GB Model Stack on an 8GB RTX 4060: LTX2.5 Generates 10s Video Locally — BigBullshitta · 2026-09-17
- 'BREATHE OUT': a creator's first AI filmmaking attempt — micheldoumit · 2026-09-17