LocateAnything Releases Open-Source Multimodal Spatial Dataset with 12M Images
ZhidingYu · x · 2026-07-28
Following the success of the LocateAnything-3B model, which reached #1 on Hugging Face Trending, the team has officially open-sourced the accompanying high-quality training dataset, LocateAnything-Data.
- Scale: Contains 12M images, 138M queries, and 785M annotations.
- Tasks Covered: Multi-object detection, referring expression grounding, physical AI, GUI understanding, OCR, and document understanding.
- Goal: To make it easier for the community to reproduce, extend, and build stronger multimodal models.
Related event: LocateAnything Open-Sources Multimodal Dataset with 12M Images(2 posts)→
More from Research
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23
- Mathematicians, not just LLMs, made AI's math breakthroughs possible, scholars argue — tak3sh8 · 2026-09-23
- AI-enabled drug discovery cuts discovery time by 15-80%, McKinsey research finds — menhguin · 2026-09-23
- Gemini training details dissected: groupwise reward redistribution to fight reward hacking — nrehiew_ · 2026-09-23