STRAP Dataset: Generating Region-Text Annotations for 2M Web Images via Frozen MLLM

ducha_aiki · x · 2026-08-07

The Visual Recognition Group at CTU Prague has open-sourced a large-scale dataset named STRAP.

It contains region-text annotations for 2 million web images (based on CC3M and DataComp). Notably, these dense captions, bounding boxes, and attributes are efficiently generated in a single forward pass using a frozen open-source Multimodal LLM (MLLM), providing rich training resources for downstream tasks like object detection and image-to-text generation.

Original post →

More from Multimodal

Multimodal channel →