EVA from UTokyo revives VAEs for sequence generation with a single extra linear layer
RichmanRonald · x · 2026-10-07
Kaede Shiohara (University of Tokyo) introduces the Empirical Variational Autoencoder (EVA), which replaces the fixed standard-Gaussian prior of VAEs with a self-predicted autoregressive latent prior via one extra linear layer on a causal decoder backbone. This yields a closer prior-posterior match and efficient ancestral sampling without VQVAE quantization or diffusion-style iterative prediction.
On ImageNet 256×256 and VGGSound, EVA achieves competitive fidelity with fewer inference parameters and faster inference versus baselines like AR-Diffusion. The recipe from VAE to EVA: make the decoder causal, reconstruct tokens while predicting the next prior, minimize KL between posterior and predicted prior, then sample ancestrally at inference.
Related event: EVA: A Single Linear Layer Brings VAEs Back to Sequence Generation(2 posts)→
More from Research
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07
- Paradigm scales RL context from 65k to 131k tokens using a trained value model — tensorqt · 2026-10-07
- Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA — tensorqt · 2026-10-07