EVA: one linear layer turns VAEs into self-prior generative models rivaling diffusion

Kaede Shiohara · hf · 2026-10-06

The authors present Empirical Variational Autoencoder (EVA), a general generative framework for continuous-valued, non-vector-quantized sequences. Built on the VAE evidence lower bound, EVA learns autoregressive latent priors empirically from data using only a single extra linear layer, replacing the standard-Gaussian prior constraint. This greatly reduces the prior-posterior distribution gap typical of VAEs and enables high-fidelity ancestral sampling. Experiments on image and sound synthesis show EVA matches autoregressive diffusion baselines in quality with much faster inference.

Related event: EVA: A Single Linear Layer Brings VAEs Back to Sequence Generation(2 posts)→

Original post →

More from Research

Research channel →