BeyondPixels: turning video diffusion latents directly into dynamic 4D scenes, skipping RGB
新智元 · wechat · 2026-08-19
BeyondPixels (ReLER/CCAI/Zhejiang Univ.) proposes Latent-to-4D: instead of decoding a video model's latents to RGB and reconstructing with a separate model, an L4AR network translates the final denoised latent directly into 4D latent, yielding dynamic 4D scenes viewable from novel angles—avoiding cross-representation error. Key insight: different video DiTs sharing the same VAE (WanVAE) speak the same "latent language", so text, image, pose and trajectory conditioning all flow through one interface. Training never runs the video DiT, uses only 1K reconstructed videos, and a single checkpoint plugs into Wan2.1-14B/1.3B and Wan2.2-I2V without retraining. In same-latent comparisons it ranks first on all Image-to-4D metrics, with 66.8%–72.1% human preference on geometry completeness. Caveats: requires shared-VAE convention, metrics aren't metric-accurate 4D, and robot cases test interface compatibility, not physics.
More from Multimodal
- InfinityEdit: Infinite Video Editing via Lightweight Adapter — Yunze Tong · 2026-08-24
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24