STRAP Dataset: Generating Region-Text Annotations for 2M Web Images via Frozen MLLM
ducha_aiki · x · 2026-08-07
The Visual Recognition Group at CTU Prague has open-sourced a large-scale dataset named STRAP.
It contains region-text annotations for 2 million web images (based on CC3M and DataComp). Notably, these dense captions, bounding boxes, and attributes are efficiently generated in a single forward pass using a frozen open-source Multimodal LLM (MLLM), providing rich training resources for downstream tasks like object detection and image-to-text generation.
More from Multimodal
- Suno Mobile App Launches Voices Feature for Custom Vocal Tracks — suno · 2026-08-07
- Seedance Workflow: Draft Cheaply with 2.0 Mini, Finish with 2.5 — Div_pradeep · 2026-08-07
- Lovart Integrates Seedance 2.5: Supports 50 Reference Inputs and 30s Videos — AIwithGhotai · 2026-08-07
- PixVerse Launches PixLight AI Film Festival with $300k Prize Pool — umesh_ai · 2026-08-07
- BytePlus Launches Seedance 2.5 Enterprise API with 30-Second Video Support — nikola_mr64990 · 2026-08-07
- AI-Generated Mashup: DOOM Meets The Shining's Overlook Hotel — ctrl-shift-face · 2026-08-07