EVA from UTokyo revives VAEs for sequence generation with a single extra linear layer

RichmanRonald · x · 2026-10-07

Kaede Shiohara (University of Tokyo) introduces the Empirical Variational Autoencoder (EVA), which replaces the fixed standard-Gaussian prior of VAEs with a self-predicted autoregressive latent prior via one extra linear layer on a causal decoder backbone. This yields a closer prior-posterior match and efficient ancestral sampling without VQVAE quantization or diffusion-style iterative prediction.

On ImageNet 256×256 and VGGSound, EVA achieves competitive fidelity with fewer inference parameters and faster inference versus baselines like AR-Diffusion. The recipe from VAE to EVA: make the decoder causal, reconstruct tokens while predicting the next prior, minimize KL between posterior and predicted prior, then sample ancestrally at inference.

Related event: EVA: A Single Linear Layer Brings VAEs Back to Sequence Generation(2 posts)→

Original post →

More from Research

Research channel →