FuseReg: Regularizing Layer Fusion to Close the Reconstruction-Generation Gap in RAEs
_akhaliq · x · 2026-09-29
- Representation Autoencoders (RAEs) fuse multi-layer encoder features into a shared latent space, but reconstruction and generation prefer different hierarchy levels: decoders favor shallow pixel-rich features while DiTs favor deeper structured semantics, creating a reconstruction–generation mismatch.
- FuseReg trains the model to stay robust across a distribution of layer fusions rather than searching for one optimal fusion, mitigating the gap.
- Paper page is live on Hugging Face.
More from Multimodal
- PrunaAI's P-Video-2 Pro models tie for #2 on Design Arena image-to-video leaderboard at Elo 1325 — guennemann · 2026-09-29
- QuiverAI's Arrow 2 Telos hits 1624 Elo, first model to break 1600 on SVG Arena leaderboard — stuffyokodraws · 2026-09-29
- Training FLUX.1 LoRAs on an 8GB RTX 5060: what optimizations work? — Wide_Director_8897 · 2026-09-29
- Flatbed debuts: an AI-native video editor where every asset is individually promptable — rchardkovacs · 2026-09-29
- One Image to a Walkable World: Hyper3D + GPT-6 Rebuilds Scenes in Three.js — Scobleizer · 2026-09-29
- Light Field Primitives: differentiable primitives replace dense ray databases for real-time novel view synthesis — zhenjun_zhao · 2026-09-29