GAE generates video and 3D geometry natively in a shared geometry latent space, code released
yshan2u · x · 2026-09-22
A team including researchers from CUHK introduced GAE, a generative model that performs true "3D-native generation" rather than reconstructing geometry from rendered RGB.
- GAE compresses geometry-foundation features into a shared latent space; a conditional flow model generates directly there.
- A single generated latent decodes both RGB appearance and 3D point clouds, aligning appearance and geometry by construction.
- Demos include camera-conditioned world generation and text-to-image pairs, with point clouds generated natively rather than reconstructed.
Code and models are already released, with a project page available.
More from Multimodal
- SenseNova U1.5 Lite beats FLUX.2-klein-9b on multi-reference image fusion in 5-task test — Kakash1i · 2026-09-22
- One-line input to finished episode: an agent pipeline built on CREAO — socialwithaayan · 2026-09-22
- AI virtual singer YURI's SURREAL closes Shanghai show on a 100-meter screen — hq4ai · 2026-09-22
- Chromovolume: visualizing human movement by stacking time into a hypnotic 3D volume — CurieuxExplorer · 2026-09-22
- Filmmaker seeks local workflow for video background replacement with matching actor lighting — ItsLukeHill · 2026-09-22
- Can Qwen Image 2.1 pair with refmods like H3? Community weighs VAE vs VLM hurdles — CranberryBeautiful76 · 2026-09-22