x0-prediction beats velocity in high-dim RAE latents, boosting text-to-image diffusion
Chao Feng · hf · 2026-09-24
- This study examines diffusability of latents in representation autoencoders (RAEs) built on pretrained visual encoders. Finetuning encoders for reconstruction recovers fine details but counterintuitively reduces effective dimensionality, altering latent geometry.
- The authors show standard velocity prediction in flow matching forces the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient.
- Clean data parameterization (x0-prediction) focuses learning on the signal manifold; experiments across multiple strong-reconstruction encoders show consistent text-to-image improvements.
Related event: ECCV 2026 paper: x0-prediction tackles high-dimensional latent diffusion(2 posts)→
More from Research
- Transformer Explainer: Interactive Visual Walkthrough of GPT-2 Internals — jackedAJ · 2026-09-24
- The coin-flip test: why LLMs fail at calibration, which is the actual product — airesearch12 · 2026-09-24
- Experts forecast AI virology parity for 2030-2034; it arrived in April 2025 — davidmanheim · 2026-09-24
- In-house multi-vec retriever crushes API dense retrievers on memory search, ColBERT next? — bo_wangbo · 2026-09-24
- OmniVChat bilingual video-dialogue dataset trends on Hugging Face — Harland · 2026-09-24
- Uranus: diffusion-based robot simulator hits 24 FPS with open-ended rollout — D-Robotics · 2026-09-24