UrbanGround benchmark: models name landmarks at 75-94% but all 10 collapse to 0-3.8% on long city navigation

jiqizhixin · x · 2026-09-15

UrbanGround, a new benchmark from Shanghai Jiao Tong University, NUS, Meituan, CUHK, Shanghai University, and Oxford, exposes a stark gap in urban spatial intelligence.

Given a street view, the best models identify landmarks at 75–93.8% accuracy. But when models must actually walk through a city, short-range navigation drops the best model to 75%, and on longer routes all 10 models fall to 0–3.8%.

Unlike prior benchmarks that stop at single street-view/aerial images or preset navigation nodes, UrbanGround builds a real 3D Hong Kong sandbox from actual geographic data where models navigate first-person: checking maps, crossing footbridges, hitting walls, encountering road closures and moving pedestrians. Answering correctly is just the starting point.

Original post →

More from Research

Research channel →