OneCanvas turns multi-view camera features into a single panoramic canvas for 2D VLMs to grasp 3D scenes

anselm · x · 2026-10-04

3D scene understanding in VLMs usually demands complex, model-specific geometry encoders or huge training budgets. OneCanvas takes a simpler route: it gathers patch features from all camera views onto a single panoramic canvas that a pretrained 2D VLM reads like an ordinary image.

Because the canvas can be centered on any pose, the same representation also supports situated reasoning from a specific viewpoint — a key capability for robotics and embodied AI.

Original post →

More from Embodied

Embodied channel →