LocateAnything-Data opens up 12M images and 785M spatial annotations for multimodal training
ZhidingYu · x · 2026-07-27
Zhiding Yu announces LocateAnything-Data, an open-source spatial dataset aimed at better multimodal models.
The release includes 12M images, 138M queries, and 785M spatial annotations, spanning detection, grounding, physical AI, GUI understanding, OCR, and document understanding. Alongside the data, the team is also open-sourcing the reproducible training pipeline: unified spatial annotations, JSONL, indexed WebDataset TARs, Megatron-Energon configs, source-media mappings, and visualization/loading examples.
Related event: LocateAnything Open-Sources Multimodal Dataset with 12M Images(2 posts)→
More from Research
- Mathematician shares a cheap 4-step heuristic for hyperparameter tuning — dejanseo · 2026-09-23
- Burkov skew AI hype: 'deterministic LLMs' and 'first agents' are old tricks rebranded — burkov · 2026-09-23
- Continuous diffusion beats discrete on random k-SAT, proposed as standard benchmark — ArashVahdat · 2026-09-23
- Grady Booch: Contemporary AI Still Lacks Abductive Reasoning, Just 'Next-Token Prediction' — Grady_Booch · 2026-09-23
- AI solves Navier-Stokes-related problem as machines upend mathematics, New Scientist reports — burny_tech · 2026-09-23
- Mathematician says OpenAI likely proved a significant partial case of the Hodge conjecture — burny_tech · 2026-09-23