OneCanvas turns a single image into 3D spatial reasoning, hitting new SOTA
Roger_M_Taylor · x · 2026-10-06
OneCanvas, a NeurIPS 2026 Spotlight paper with fully open-source code and model, merges video frames of a room into one panorama so an ordinary image AI can answer questions about 3D space.
- New state of the art on spatial reasoning benchmarks
- Answers from any viewpoint, including the AI's own position in the room
- Requires depth data and camera positions, not just plain video
More from Multimodal
- Gaussian GRPO normalizes multimodal RL reward distributions, boosting OpenVLThinker v2 — kaiwei_chang · 2026-10-06
- Dev uses GitHub Copilot app to generate a rebuildable 60s hype video via PR — DanWahlin · 2026-10-06
- Tencent releases new open-weight video generation model — Famous-Sport7862 · 2026-10-06
- Nano Banana 2.1 on Google Flow stuns with editorial portrait quality — full prompt shared — aziz4ai · 2026-10-06
- TagScribeR rebuilt: free local dataset studio with native LoRA training on AMD ROCm and NVIDIA — ArchAngelAries · 2026-10-06
- ComfyUI queue stuck? A maintainer's checklist to separate validation, node failures and lost progress — fluxdraw · 2026-10-06