CVP from UC San Diego and Lambda beats Video-3D-LLM on all five 3D spatial reasoning benchmarks
TheZachMueller · x · 2026-09-25
Lambda's blog details CVP (UC San Diego + Lambda, WACV 2026), a method that fixes 3D vision-language models misidentifying task-relevant objects (e.g., guessing "sewing machine" instead of a tray rack).
Key ideas:
- A target-affinity token highlights task-relevant objects.
- An allocentric grid encodes global spatial context, with a central/peripheral token split.
Against Video-3D-LLM: SQA3D EM improves 58.6 → 62.3 and Scan2Cap CIDEr 83.8 → 90.5, with gains across all five benchmarks (ScanQA, SQA3D, ScanRefer, Multi3DRefer, Scan2Cap). The post frames spatial reasoning — understanding 3D structure and object relations, not just 2D perception — as essential for physical AI and robotics.
More from Research
- Researcher lands 4 NeurIPS papers including one Oral, spanning synthesis to protein diffusion — abeirami · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Yoav Goldberg: LLM reasoning traces are 'too good' — unclear how they emerge from RL — yoavgo · 2026-09-25
- NeurIPS Oral: IDS co-evolves code with formal proofs, hitting 3x Claude Code's success rate — adityagp · 2026-09-25
- Neuralese Recurrence Explained: Why OpenAI's Astra Loop Is Not Yet Unreadable AI Thought — Astral Codex Ten · 2026-09-25
- AI Agents Are Noisy Representatives: Paper Finds Buyers End Up With Their 5th Choice Out of 10 — ghadfield · 2026-09-25