Sphere Encoder 2: Turning an Autoencoder into a 1-4 Step Image Generator

kastnerkyle · x · 2026-10-03

Tom Goldstein's group released Sphere Encoder 2 (arXiv:2610.02208), which turns an autoencoder into a fast standalone image generator. The paper identifies two flaws in the original formulation: random latent points concentrate near the equator of the encoded sphere while the training rotation never reaches that region, leaving a gap that limits one-step generation; and pixel-wise reconstruction loss during generation training pushes the decoder to average over plausible images, producing blurry outputs lacking high-frequency detail.

Sphere Encoder 2 fixes both by fully covering the spherical latent space, separating reconstruction from generation, and adding semantic alignment plus score matching. It substantially improves generation quality while keeping the speed and simplicity of an autoencoder, achieving strong ImageNet results in just 1-4 steps. Models are released and code is coming.

Original post →

More from Multimodal

Multimodal channel →