OneCanvas: Efficient 3D Scene Understanding via Panoramic Reprojection
rsasaki0109 · x · 2026-08-21
OneCanvas introduces a novel approach to 3D scene understanding for Vision-Language Models (VLMs) that avoids complex geometry encoders or high training costs. It aggregates patch features from multiple views onto a single equirectangular panoramic canvas. By unprojecting patches to 3D world coordinates using depth and pose, and adding 3D positional embeddings, OneCanvas maintains spatial context without rasterization, enabling efficient spatial reasoning across a unified coordinate system.
More from Research
- Chain-of-Experience Enables Continual LLM Improvement with Lower Cost — UCSC-VLAA · 2026-08-21
- New SONIC robot checkpoint adds finer-grained wrist manipulation, works best with PICO 5-sensor mode — zhengyiluo · 2026-08-21
- CUHK Team Open Sources Libra: 3x Throughput for Agentic Training — jiqizhixin · 2026-08-21
- Deep Dive: How Prompts, Params, and Engines Skew LLM Benchmarks — rsasaki0109 · 2026-08-21
- CMU et al. release DelusionEval, revealing LLMs reinforce delusions and safety failures grow with conversation length — burkov · 2026-08-21
- Pretraining Potential: Coding Agents and the Compute Bottleneck — zeeshanp_ · 2026-08-21