UrbanGround benchmark: models name landmarks at 75-94% but all 10 collapse to 0-3.8% on long city navigation
jiqizhixin · x · 2026-09-15
UrbanGround, a new benchmark from Shanghai Jiao Tong University, NUS, Meituan, CUHK, Shanghai University, and Oxford, exposes a stark gap in urban spatial intelligence.
Given a street view, the best models identify landmarks at 75–93.8% accuracy. But when models must actually walk through a city, short-range navigation drops the best model to 75%, and on longer routes all 10 models fall to 0–3.8%.
Unlike prior benchmarks that stop at single street-view/aerial images or preset navigation nodes, UrbanGround builds a real 3D Hong Kong sandbox from actual geographic data where models navigate first-person: checking maps, crossing footbridges, hitting walls, encountering road closures and moving pedestrians. Answering correctly is just the starting point.
More from Research
- Voodoo Dynamic Quant Goes MIT: Gradient Descent Picks Per-Tensor Quant Levels — 1ncehost · 2026-09-15
- Aphantasia and abstract symbols: new lens on the evolution of human language and culture — abenitezburraco · 2026-09-15
- Hybrid Route: LLMs Synthesize Probabilistic Programs for Sound Reasoning, as Shown in Tenenbaum-Linked Paper — xuanalogue · 2026-09-15
- Probabilistic Programming Defended: Learning Yields Compact Symbolic Program Libraries, Not Numeric Circuits — xuanalogue · 2026-09-15
- FlockMTL DuckDB extension brings LLM and RAG functions directly into SQL — _reachsumit · 2026-09-15
- Tencent Hunyuan's EvolveScaler benchmark drops frontier models to 11.3 median on hardest tier — TencentHunyuan · 2026-09-15