Qwen-3D: Enhancing Spatial Reasoning with 3D Geometry

Qwen-3D is a general 3D vision-language model submitted to ECCV 2026 that leverages depth and camera poses to fuse multi-view inputs in a shared 3D space. By compressing visual information through 3D geometry, it overcomes long-video reasoning bottlenecks and simultaneously handles grounding, segmentation, VQA, and spatial reasoning.

2026-08-07 ~ 2026-08-08 · 2 related posts