Empirical Variational
Autoencoder

A simple empirical latent prior turns VAEs into high-fidelity generators for continuous-valued sequences.

Kaede Shiohara
The University of Tokyo

Curated image samples generated by EVA-L on ImageNet
Overview

Learn the prior from the data.

EVA replaces the fixed standard-Gaussian constraint in conventional VAEs with a self-predicted autoregressive latent prior. The result is a closer match between prior and posterior, yielding efficient ancestral sampling without quantization (VQVAE) or iterative denoising (diffusion).

How it works

One linear layer revives VAEs.

EVA keeps the familiar encoder–decoder structure of VAEs and learns the autoregressive prior with an additional linear layer for prior prediction on the causal decoder backbone.

EVA framework for training and inference
4 steps from VAE to EVA generation.
01Make the VAE decoder "causal".
02Reconstruct the input token and predict the next prior.
03Minimize KL divergence between the posterior and the predicted prior.
04Sample latents ancestrally and decode them causally at inference.
ImageNet-256 & VGGSound

Competitive fidelity, efficient sampling.

Quantitative ImageNet and VGGSound results for EVA and baselines

Comparison with Baselines on ImageNet (256×256) and VGGSound. AR-Diffusion (12 blocks) has the same number of Transformer blocks as EVA at inference time while AR-Diffusion (24 blocks) has the same number of Transformer blocks as EVA at training time. In the right, we visualize AR-Diffusion with different denoising steps {10, 15, 20, 30, 40, 50}. Despite the fewer inference parameters and faster inference time, our method achieves competitive generation fidelity.

VGGSound samples

Hear EVA generate.

Each clip contains a real reference and three independently generated samples. Select a variant, then play it—only one clip plays at a time.

Latent-space visualizations for Causal VAE and EVA
Each patch represents bi-directional KL divergence between the source latent (+) and each latent. Lighter indicates similar to the source patch.
Framework analysis

Latent Space Visualization

In EVA’s latent space, the semantically similar patches tend to have similar latent distributions, implying that learning generation naturally encourages the model to learn semantics.

Citation

Use EVA in your work.

Please cite the arXiv version when using EVA in your work.

View on arXiv
@article{shiohara2026empirical,
  title={Empirical Variational Autoencoder},
  author={Kaede Shiohara},
  journal={arXiv:2610.06545},
  year={2026}
}