BeyondPixels: turning video diffusion latents directly into dynamic 4D scenes, skipping RGB

新智元 · wechat · 2026-08-19

BeyondPixels (ReLER/CCAI/Zhejiang Univ.) proposes Latent-to-4D: instead of decoding a video model's latents to RGB and reconstructing with a separate model, an L4AR network translates the final denoised latent directly into 4D latent, yielding dynamic 4D scenes viewable from novel angles—avoiding cross-representation error. Key insight: different video DiTs sharing the same VAE (WanVAE) speak the same "latent language", so text, image, pose and trajectory conditioning all flow through one interface. Training never runs the video DiT, uses only 1K reconstructed videos, and a single checkpoint plugs into Wan2.1-14B/1.3B and Wan2.2-I2V without retraining. In same-latent comparisons it ranks first on all Image-to-4D metrics, with 66.8%–72.1% human preference on geometry completeness. Caveats: requires shared-VAE convention, metrics aren't metric-accurate 4D, and robot cases test interface compatibility, not physics.

Original post →

More from Multimodal

Multimodal channel →