LocateAnything-Data opens up 12M images and 785M spatial annotations for multimodal training
ZhidingYu · x · 2026-07-27
Zhiding Yu announces LocateAnything-Data, an open-source spatial dataset aimed at better multimodal models.
The release includes 12M images, 138M queries, and 785M spatial annotations, spanning detection, grounding, physical AI, GUI understanding, OCR, and document understanding. Alongside the data, the team is also open-sourcing the reproducible training pipeline: unified spatial annotations, JSONL, indexed WebDataset TARs, Megatron-Energon configs, source-media mappings, and visualization/loading examples.
More from Research
- ReactBench v1 ranks GPT-5.6 at about 53% and Opus 5 at 49% on realistic React tasks — aidenybai · 2026-07-28
- fchollet on iLands: Open-Ended Survival Environments Are Key to Shaping AI Intelligence — fchollet · 2026-07-27
- NVIDIA releases Molt, a PyTorch-native framework for agentic RL — dair_ai · 2026-07-27
- ICML 2025 paper shows denoising models can raise spatial reasoning accuracy from under 1% to over 50% — CSProfKGD · 2026-07-27
- AI One-Shots 25k Lines of Code to Simulate Black Hole Electrodynamics in 5 Minutes — kylekabasares · 2026-07-27
- AutoScientist challenge asks systems to optimize training toward any objective — sarahookr · 2026-07-27