Michigan Researchers Distill T5Gemma-2 Embeddings Into a More Diffusible Latent Space, Beating GPT-2-M

umich · hf · 2026-10-02

A University of Michigan study asks which embedding makes the best latent space for continuous diffusion language models. Scaling the embedding model (T5 → T5Gemma-1 → T5Gemma-2) improves generation, but raw T5Gemma-2 embeddings are overly discriminative, so diffusion often lands on invalid embeddings. Distilling T5Gemma-2 into a student encoder using teacher decoded probabilities as soft labels yields a more connected latent space. The resulting mid-sized DLM reaches Gen. PPL 17.8 on OpenWebText (vs. real-text PPL 15.4), outperforming GPT-2-M.

Original post →

More from Research

Research channel →