Google Intrinsic's 3PT unifies detection, segmentation and 6DoF pose, wins both BOP 2025 tracks
kscottz · x · 2026-08-27
3PT (3D-Object Perception Transformer) from Google Intrinsic unifies detection, segmentation, and 6DoF pose estimation into two RGB-only multi-view transformers. It placed first by significant margins in both the Industrial Robotics and AR/VR tracks of the BOP 2025 challenge at ICCV, and the paper is a CVPR 2026 Highlight.
The model demonstrates strong cross-domain robustness and is already deployed in real-world industrial robotic workcells worldwide as the Intrinsic Vision Model (IVM).
OpenCV Live episode 222 features paper author Agastya Kalra (Intrinsic, University of Hawaii), walking through the 3PT architecture, BOP 2025 results, and what it took to move the model from benchmark to production robotics.
More from Embodied
- QNX partners with Hailo for edge Physical AI: 14x performance consistency — pdamodaran · 2026-08-27
- Robotics Moat: Data, Hardware, and the Deployment Layer Flywheel — chris_j_paxton · 2026-08-27
- End-to-End RL Drone Policy Passes Sim2real on Multiple Hardware — yacineMTB · 2026-08-27
- Robot that looks like a Sony Walkman and can roller skate — Thom_Wolf · 2026-08-27
- Waymo, Zoox Test Drivers Injured by Robotaxi Sudden Braking — tbirdcymru · 2026-08-27
- Why LLMs Can't Count Windows: Elorian CEO on Multimodal Reasoning — bigdata · 2026-08-27