Fourth Perception Test Challenge at ECCV 2026 pushes multimodal models on city-scale spatial intelligence
AjdDavison · x · 2026-09-09
The 4th Perception Test Challenge was held at ECCV 2026 in Malmö on Sept 9, themed 'Spatial Intelligence: from Table-top to City-scale', with €20K in prizes testing whether multimodal models truly grasp spatial structure.
- Tracks: Unified VQA, Grounded VQA, plus new KilometerVision (distance estimation, landmark recognition, compass and map understanding in hour-long walking videos) and KilometerAudio (audio-video multimodal understanding).
- Keynote: UC Berkeley's Jack L. Gallant on human navigation — VR city training plus fMRI taxi tasks across 38 feature spaces found naturalistic navigation supported by 11 distinct regions spanning visual, parietal and prefrontal cortices.
- DeepMind's David Davison spoke at the main-arena workshop.
More from Multimodal
- Marigold V2 retools diffusion transformers for sharper monocular depth estimation — huawei-bayerlab · 2026-09-09
- One-Take Rainy Toll Booth: A Full AI Video Prompt for Korean Supernatural Mystery — umesh_ai · 2026-09-09
- MiniMax H3 Ref2V resists surgical video edits, users hunt for prompting workarounds — Weird_Ad4978 · 2026-09-09
- Full AI video pipeline: Qwen 3.8 + Flux 2 Klein + Minimax H3 + LTX 2.5 + Breeze TTS — CQDSN · 2026-09-09
- Minimax H3 videos look h264-compressed regardless of settings, Redditor reports — FoxTrotte · 2026-09-09
- Mage brings MiniMax H3 & H3 Turbo with unlimited video generation, LoRA support — MiniMax_AI · 2026-09-09