FuseReg: Layer-Fusion Regularization Cuts gFID up to 29% in Representation Autoencoders
USC-PSI-Lab · hf · 2026-09-28
USC PSI Lab proposes FuseReg, tackling the layer-selection trade-off in representation autoencoders (RAEs) for image generation.
- Problem: shallow encoder layers preserve pixel detail while deep layers yield better generation metrics; fixed heuristic fusion couples two stages wanting different information
- Method: train over random subsets of encoder layers; theoretically, subset sampling explicitly penalizes sensitivity to cross-layer disagreement
- Results (ImageNet-256, DINOv3-L): one decoder reconstructs from full, sparse, and single-layer fusions without retraining at higher PSNR than fusion-specialized decoders; decoder replacement alone cuts unguided gFID by 27% on an unchanged RAEv2 DiT-XL generator; extending the regularization to diffusion training cuts gFID by a further 29% on DiT-Base
Takeaway: training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without touching the pretrained encoder.
More from Research
- InternW0-Δ open-sources code, weights and 20K+ hours of robot data — arankomatsuzaki · 2026-09-28
- IMLE-VLA replaces flow matching with single-step cIMLE, 3.67x faster robot actions at 55Hz — petitegeek · 2026-09-28
- Red Queen Bio Uses AI and Wet Labs to Build Medicines Against Natural and Synthetic Viruses — _sholtodouglas · 2026-09-28
- Yacine demos general sim2real pipeline: policy trained in just two minutes — yacineMTB · 2026-09-28
- Claude Spots a Missing 1/2 Coefficient in a Physics Preprint — tak3sh8 · 2026-09-28
- Astra saw through perturbed Zork observations, calling it a 'variant' — tw_killian · 2026-09-28