CVP from UC San Diego and Lambda beats Video-3D-LLM on all five 3D spatial reasoning benchmarks

TheZachMueller · x · 2026-09-25

Lambda's blog details CVP (UC San Diego + Lambda, WACV 2026), a method that fixes 3D vision-language models misidentifying task-relevant objects (e.g., guessing "sewing machine" instead of a tray rack).

Key ideas:

Against Video-3D-LLM: SQA3D EM improves 58.6 → 62.3 and Scan2Cap CIDEr 83.8 → 90.5, with gains across all five benchmarks (ScanQA, SQA3D, ScanRefer, Multi3DRefer, Scan2Cap). The post frames spatial reasoning — understanding 3D structure and object relations, not just 2D perception — as essential for physical AI and robotics.

Original post →

More from Research

Research channel →