Object-Uni Turns Object Pose Into a Shared Variable for Understanding and Generation, With One-Third of GPT-4o's Azimuth…
青稞AI · wechat · 2026-10-10
The Institute of Automation of the Chinese Academy of Sciences, the School of Artificial Intelligence of the University of Chinese Academy of Sciences, Tsinghua University, and Ant Group jointly released a paper on Object-Uni, a unified model for object-level spatial understanding and controllable generation. Its core idea is to treat object pose as an explicit geometric variable shared between understanding and generation, enabling orientation estimation, spatial relation reasoning, pose-controllable text-to-image generation, and object-level novel view synthesis within a single framework.
More from Multimodal
- Step 5 Preview generates a 30-second motion clip in one shot via Hermes agent — Teknium · 2026-10-11
- Tencent's open-source Hunyuan3D-2 turns one image into textured 3D models, 15k stars — Promptmethus · 2026-10-11
- Open-source music model YuE2 runs locally on a MacBook Pro, rivaling Suno quality — vista8 · 2026-10-11
- Two prompts + Claude made a Monet-style MV for Jay Chou's 'Qi Li Xiang' — AlchainHust · 2026-10-11
- Underdog launches on-device image gen powered by Qwen models, photos never leave your computer — Scobleizer · 2026-10-11
- Open-source anime DiT models: Anima staleness has users pinning hopes on Krea2 fine-tunes — Pristine_Stress_670 · 2026-10-11