Object-Uni Turns Object Pose Into a Shared Variable for Understanding and Generation, With One-Third of GPT-4o's Azimuth…

青稞AI · wechat · 2026-10-10

The Institute of Automation of the Chinese Academy of Sciences, the School of Artificial Intelligence of the University of Chinese Academy of Sciences, Tsinghua University, and Ant Group jointly released a paper on Object-Uni, a unified model for object-level spatial understanding and controllable generation. Its core idea is to treat object pose as an explicit geometric variable shared between understanding and generation, enabling orientation estimation, spatial relation reasoning, pose-controllable text-to-image generation, and object-level novel view synthesis within a single framework.

Original post →

More from Multimodal

Multimodal channel →