UrbanGround benchmark: LLMs' long-range city navigation drops to 0%
机器之心 · wechat · 2026-09-05
Researchers from SJTU, NUS, Meituan, CUHK and Oxford built UrbanGround, a navigable 3D city sandbox from real Hong Kong geodata, with 810 verified tasks testing 10 multimodal models.
- Visual recognition reaches 75.0–93.8% accuracy, but direction understanding only 23.3–58.3% (Gemini-3.1-Pro below random).
- Short-range navigation peaks at 75.0%; long-range collapses to 0–3.8% (GPT-5.5: 75%→0%).
- With road closures and pedestrians, compliance drops to 10.0–46.7% and collision rates hit 76.3–90.0%.
The takeaway: models "forget the road after a few steps" — each step's judgment fails to carry through. Paper, code and a playable app are public.
More from Multimodal
- Creator's Seedance AI video hailed as setting a new bar for quality — LinusEkenstam · 2026-09-05
- Astra builds an accurate 3D model of Chambord castle and renders it cinematically in Blender — Dimillian · 2026-09-05
- GPT-6 Astra powers 3D website that explodes a Tesla Model X into 334 modeled parts — Dimillian · 2026-09-05
- Seedance 2.5 Generates a Realistic 30-Second Zombie Bus Horror Film From One Prompt — SimplyAnnisa · 2026-09-05
- Two 45-minute loops with a critic: how GPT-6 Astra iterates on 3D reconstruction — bilawalsidhu · 2026-09-05
- GPT-6 Astra rebuilds a living room in Blender from a 3D scan, no downloaded assets — bilawalsidhu · 2026-09-05