PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

zhenjun_zhao · x · 2026-07-03

The study proposes PointDiT, a monocular geometry estimation method based on pixel-space diffusion. Conditioned on DINOv3, the model uses a ViT architecture to directly output point map patches, predicting 3D geometric structures from a single image.

Compared to traditional depth estimation methods, this solution introduces the generative power of diffusion models into 3D point cloud reconstruction. It can recover dense geometric information from standard RGB images, offering potential value for downstream applications like autonomous driving, AR/VR, and 3D reconstruction.

Related event: PointDiT: A Minimalist Point-Space Diffusion Method for 3D Reconstruction(2 posts)→

Original post →

More from Multimodal

Multimodal channel →