Tencent ARC's GAE: geometry-native latent space cuts FVD by up to 23% in world generation
TencentARC · hf · 2026-09-23
Tencent ARC presents GAE (geometry-native autoencoder), arguing inconsistent 3D generation is a representation problem, not just modeling: generators evolve appearance-centric latents while perception models recover geometry in cross-view-aware space. GAE reparameterizes a geometry foundation model's features into a compact latent jointly decodable to appearance, depth, cameras, and point maps; a standard conditional flow then supports diverse generation tasks.
In controlled comparisons holding the generator and training fixed, swapping in the GAE latent improves both visual quality and independently measured 3D coherence: FVD falls 12.7% and 23.1% on RealEstate10K and DL3DV, and camera-trajectory error halves on RealEstate10K — showing the latent space is central to geometry-consistent generation and a shared perception-generation interface.
More from Multimodal
- Gradio shrinks Qwen-Image 2.1's 9B prompt rewriter to 0.8B that runs on laptops — Gradio · 2026-09-23
- Mirage launches Tesseract, a video creative suite that lets AI agents edit footage directly — ccerrato147 · 2026-09-23
- Opus 5.5 one-shots a 7-minute tutorial video, Reddit user stunned — thatisnotmychapstick · 2026-09-23
- Flux 3 aces split-screen rendering, keeping two camera angles mostly in sync — umesh_ai · 2026-09-23
- RULER: instance-aware rubric rewards beat reward hacking in SVG generation RL — inclusionAI · 2026-09-23
- PixVerse unveils R2 real-time world model: actions carry cause and effect — alifcoder · 2026-09-23