Evaluating Multimodal Models on Geospatial Image Localization
Afinetheorem · x · 2026-07-10
The author evaluated several models on a global geospatial image localization benchmark similar to Geoguessr. The test is exceptionally difficult, as most images lack obvious clues like text or street signs.
Results show that 70% of images have already been pinpointed exactly by a model, while 86% are located within a 100-mile error radius. Model performances vary significantly: Terra reached the level of Gemini 2.0 Flash, whereas Luna only matched Sonnet 3.5/Grok 4.5, making it one of the worst-performing OpenAI models in the test.
Related event: EyeBench-v3 Updates: Sol Model Takes the Lead(4 posts)→
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22
- Gemini 3.5 Flash Lite Tested: Not Frontier-Optimal, but Hits 350 tok/s — brandon_galang · 2026-07-22