DRS-VPT: feed-forward camera pose estimation from a point cloud scan and a single image

kwangmoo_yi · x · 2026-09-16

Fu and Fallon introduce DRS-VPT, a vision point transformer that performs feed-forward camera pose estimation given a colorless point cloud scan and an image, built on DINO + Sonata + DPT with a scale component.

Related event: Oxford Researchers Unveil DRS-VPT for Direct Image-to-Point-Cloud Relocalization(2 posts)→

Original post →

More from Embodied

Embodied channel →