Paper PixWorld: Unifying 3D Generation and Reconstruction in Pixel Space

pmttyji · reddit · 2026-07-07

The paper PixWorld unifies 3D reconstruction and generation into a single pixel-space diffusion paradigm, using one model to tackle both tasks. It supervises diffusion directly on rendered images without needing pre-trained VAE/RAE, and introduces geometry-aware loss to align within the 3D foundation model's feature space. It surpasses previous latent-space methods in generation and matches SOTA in reconstruction.

Related event: PixWorld Unifies 3D Scene Generation and Reconstruction(3 posts)→

Original post →

More from Multimodal

Multimodal channel →