ECCV 2026 paper: x0-prediction fixes inefficient diffusion in reconstruction-tuned RAE latent spaces
serrjoa · x · 2026-09-24
A paper accepted to ECCV 2026 studies the diffusibility of latents in Representation AutoEncoders (RAEs).
Key findings
- Off-the-shelf visual encoders discard fine-grained details; finetuning for reconstruction recovers them but reduces effective latent dimensionality, altering the geometry.
- Standard velocity prediction in flow matching then has to fit orthogonal noise directions off the low-dimensional signal manifold, making optimization inefficient.
- Switching to x0-prediction (clean data parameterization) focuses learning on the signal manifold.
Conclusion: across multiple strong-reconstruction encoders, x0-prediction consistently improves text-to-image generation without latent compression. Authors include Chao Feng, Yijun Li, and Richard Zhang.
Related event: ECCV 2026 paper: x0-prediction tackles high-dimensional latent diffusion(2 posts)→
More from Multimodal
- Qwen Image 2.1 falls back to image editing when given reference images — extra2AB · 2026-09-24
- Kling 4.0 leak: 120-second videos, Omni Reference with 15 elements, up to 4K images — koltregaskes · 2026-09-24
- Limestone claims Claude Opus 5.5 one-shotted its entire launch video — alex_verem · 2026-09-24
- Using Ling-3.0-flash-VL to analyze 8 ultrasound reports and solve a pregnancy puzzle — alifcoder · 2026-09-24
- Claude synthesizes a full band from math, noise and its own voice — willcb · 2026-09-24
- Viggle's Qwen-Image-2.1-viggle-turbo hits v0.2.1 amid quality complaints — reeight · 2026-09-24