Michigan Researchers Distill T5Gemma-2 Embeddings Into a More Diffusible Latent Space, Beating GPT-2-M
umich · hf · 2026-10-02
A University of Michigan study asks which embedding makes the best latent space for continuous diffusion language models. Scaling the embedding model (T5 → T5Gemma-1 → T5Gemma-2) improves generation, but raw T5Gemma-2 embeddings are overly discriminative, so diffusion often lands on invalid embeddings. Distilling T5Gemma-2 into a student encoder using teacher decoded probabilities as soft labels yields a more connected latent space. The resulting mid-sized DLM reaches Gen. PPL 17.8 on OpenWebText (vs. real-text PPL 15.4), outperforming GPT-2-M.
More from Research
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence — Haibo Wang · 2026-10-02
- Netflix's Align Then Reason lip-sync judge boosts mean AUC by up to 59% — netflix · 2026-10-02
- Peking University's DexPolicy lifts dexterous manipulation success to 85% via annealed exploration — PekingUniversity · 2026-10-02
- Microsoft's ActiveSaddler Uses Automated Curriculum Learning to Boost Agent Harnesses by 7.5 Points — microsoft · 2026-10-02
- Alibaba's PoS Maintains Explicit Belief States to Fix Long-Horizon Agent 'Belief Trapping' — alibabagroup · 2026-10-02