Google Intrinsic's 3PT unifies detection, segmentation and 6DoF pose, wins both BOP 2025 tracks

kscottz · x · 2026-08-27

3PT (3D-Object Perception Transformer) from Google Intrinsic unifies detection, segmentation, and 6DoF pose estimation into two RGB-only multi-view transformers. It placed first by significant margins in both the Industrial Robotics and AR/VR tracks of the BOP 2025 challenge at ICCV, and the paper is a CVPR 2026 Highlight.

The model demonstrates strong cross-domain robustness and is already deployed in real-world industrial robotic workcells worldwide as the Intrinsic Vision Model (IVM).

OpenCV Live episode 222 features paper author Agastya Kalra (Intrinsic, University of Hawaii), walking through the 3PT architecture, BOP 2025 results, and what it took to move the model from benchmark to production robotics.

Original post →

More from Embodied

Embodied channel →