PointDiT: Pixel-Space Diffusion Transformer for Monocular Geometry

google · hf · 2026-07-08

Google proposed PointDiT, which employs a pure ViT architecture to directly process 3D point map patches, combined with DINOv3 image tokens for conditional generation. While maintaining a minimalist architecture, this method surpasses complex latent-space models and demonstrates better robustness in ambiguous regions.

Related event: Google's PointDiT Recovers 3D Shapes from Single Images(2 posts)→

Original post →

More from Research

Research channel →