Testing Multimodal LLMs on GeoGuessr: Impressive but Slower Than Humans
HanchungLee · x · 2026-08-04
A developer tested several multimodal LLMs on a GeoGuessr task. Earlier attempts by Gemini and other models failed hard after being misled by the word 'cupertino'. A recent test showed that a 5.6 version model solved the task in 10 minutes using two geolocates and browser control, demonstrating strong multimodal reasoning and tool use, though it remains much slower than human experts who can do it in under 3 minutes.
Related event: Multimodal LLMs Struggle with Geo-Guessing Tasks(2 posts)→
More from Models
- Analyzing Kimi-K3 Open Source: Data Flywheels Are the New Moat — karminski3 · 2026-08-04
- US Frontier Labs Tout 'Efficiency' as Chinese Models Drop at Throwaway Prices — deliprao · 2026-08-04
- Qwen 3.8 Coding Test: Nearly Matches K3 at Half the Price — bindureddy · 2026-08-04
- Fable and Gemini Vision Models Misled by Place Names in Geo-Guessing Fail — HanchungLee · 2026-08-04
- Opinion: Masked Language Modeling Was a Detour, Autoregressive Was Inevitable — jxmnop · 2026-08-04
- DeepMind Releases DiffusionGemma: Discrete Diffusion for Ultra-Fast Text Generation — deepmind · 2026-08-04