GeoGuess Bench: Claude Opus 5.5 tops 210-photo geolocation test, open models shine
nutlope · x · 2026-10-07
Nutlope launched GeoGuess Bench, a benchmark measuring how well AI models play GeoGuessr: each model gets 210 photos from around the world, predicts where each was taken, and is scored on proximity.
Surprising results:
- Muse Glimmer 30B beat GPT 6 Astra while being 45x cheaper.
- GLM 5.3 Flash matched GPT 6.1 Sol at 5x cheaper.
- Claude Opus 5.5 took #1, scoring above Fable 5.1.
Open models did surprisingly well overall. The project includes a full leaderboard, Pareto frontier curve, and every model's guesses, with code open-sourced.
More from Models
- Google's multimodal embeddinggemma-2 trends on Hugging Face — google · 2026-10-07
- Mistral CEO: Large 4 trained on our own compute, 'RL shows no sign of saturation' — sivareddyg · 2026-10-07
- Perplexity ships open-weights pplx-decider-v1.1-27b at half the cost of v1 — perplexity_ai · 2026-10-07
- Mistral launches Large 4: 1T-param multimodal model, 49B active, open weights in October — beffjezos · 2026-10-07
- Perplexity's open-weights pplx-decider-v1.1-27b tops Hugging Face Decision Index 0.3 — AravSrinivas · 2026-10-07
- AutoAWQ Author: Reproduce Bonsai 2-Class Ternary Model for ~$43k on One B300 Node in ~4 Weeks — airesearch12 · 2026-10-07