OneCanvas turns multi-view camera features into a single panoramic canvas for 2D VLMs to grasp 3D scenes
anselm · x · 2026-10-04
3D scene understanding in VLMs usually demands complex, model-specific geometry encoders or huge training budgets. OneCanvas takes a simpler route: it gathers patch features from all camera views onto a single panoramic canvas that a pretrained 2D VLM reads like an ordinary image.
Because the canvas can be centered on any pose, the same representation also supports situated reasoning from a specific viewpoint — a key capability for robotics and embodied AI.
More from Embodied
- This Robot Digs Through Solid Rock Without Ever Touching It — TinfoilTricorn · 2026-10-04
- Dev builds a physical control deck to juggle multiple AI coding agent sessions — Weak_Cherry_4878 · 2026-10-04
- Frontier VLMs Struggle at Basic 2-DoF Active Visual Search, NeurIPS Paper Finds — _vztu · 2026-10-04
- NYU's $5 open-source eFlesh magnetic tactile sensor is 3D-printable — lukas_m_ziegler · 2026-10-04
- SPEAR Simulator Opens Up 14K+ Unreal Engine Functions for Embodied AI Research — rsasaki0109 · 2026-10-04
- vLLM creator's new 'roommate' walks like he's had three drinks and says no to everything — Thom_Wolf · 2026-10-04