LocateAnything-Data opens up 12M images and 785M spatial annotations for multimodal training

ZhidingYu · x · 2026-07-27

Zhiding Yu announces LocateAnything-Data, an open-source spatial dataset aimed at better multimodal models.

The release includes 12M images, 138M queries, and 785M spatial annotations, spanning detection, grounding, physical AI, GUI understanding, OCR, and document understanding. Alongside the data, the team is also open-sourcing the reproducible training pipeline: unified spatial annotations, JSONL, indexed WebDataset TARs, Megatron-Energon configs, source-media mappings, and visualization/loading examples.

Original post →

More from Research

Research channel →