Evaluating Multimodal Models on Geospatial Image Localization

Afinetheorem · x · 2026-07-10

The author evaluated several models on a global geospatial image localization benchmark similar to Geoguessr. The test is exceptionally difficult, as most images lack obvious clues like text or street signs.

Results show that 70% of images have already been pinpointed exactly by a model, while 86% are located within a 100-mile error radius. Model performances vary significantly: Terra reached the level of Gemini 2.0 Flash, whereas Luna only matched Sonnet 3.5/Grok 4.5, making it one of the worst-performing OpenAI models in the test.

Related event: EyeBench-v3 Updates: Sol Model Takes the Lead(4 posts)→

Original post →

More from Models

Models channel →