Tencent ARC's GAE: geometry-native latent space cuts FVD by up to 23% in world generation

TencentARC · hf · 2026-09-23

Tencent ARC presents GAE (geometry-native autoencoder), arguing inconsistent 3D generation is a representation problem, not just modeling: generators evolve appearance-centric latents while perception models recover geometry in cross-view-aware space. GAE reparameterizes a geometry foundation model's features into a compact latent jointly decodable to appearance, depth, cameras, and point maps; a standard conditional flow then supports diverse generation tasks.

In controlled comparisons holding the generator and training fixed, swapping in the GAE latent improves both visual quality and independently measured 3D coherence: FVD falls 12.7% and 23.1% on RealEstate10K and DL3DV, and camera-trajectory error halves on RealEstate10K — showing the latent space is central to geometry-consistent generation and a shared perception-generation interface.

Related event: Tencent ARC's GAE: Geometry-Native Latent Space for 3D-Consistent Generation(3 posts)→

Original post →

More from Multimodal

Multimodal channel →