Testing Multimodal LLMs on GeoGuessr: Impressive but Slower Than Humans

HanchungLee · x · 2026-08-04

A developer tested several multimodal LLMs on a GeoGuessr task. Earlier attempts by Gemini and other models failed hard after being misled by the word 'cupertino'. A recent test showed that a 5.6 version model solved the task in 10 minutes using two geolocates and browser control, demonstrating strong multimodal reasoning and tool use, though it remains much slower than human experts who can do it in under 3 minutes.

Related event: Multimodal LLMs Struggle with Geo-Guessing Tasks(2 posts)→

Original post →

More from Models

Models channel →