Uno adds diffusion weights to AR LLMs, beating EAGLE-3 and all diffusion LLMs
JFPuget · x · 2026-09-04
New method Uno tackles diffusion LLMs' two key limitations vs AR models: lower quality and slower inference at large batch sizes.
- Keeps the AR architecture of LLMs, with each layer holding two weight sets: AR weights and diffusion weights
- The diffusion weights enable lossless parallel sampling from the AR distribution
- Faster than all speculative decoding methods, including DFlash and EAGLE-3
- Outperforms every existing diffusion LLM: Mercury 2, Diffusion Gemma, and Llada
Paper, models, and code are all linked.
Related event: Uno: diffusion-augmented LLM matches AR quality with faster inference(2 posts)→
More from Models
- Critics warn OpenAI's GPT-6 Astra reasons opaquely, gutting CoT monitoring safety — GaryMarcus · 2026-09-04
- ChatGPT adds writing-style matching from connected apps, analytics, and a Yubikey deal tied to Daybreak access — btibor91 · 2026-09-04
- GPT-6 reportedly launches as Tesla starts public rides in steering-free Cybercab — Dr_Singularity · 2026-09-04
- Astra early-access users' similar blender demos look coordinated, with no practical examples shown — jdjohnson · 2026-09-04
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04
- OpenAI engineer says GPT-6 "Astra" has "tremendously" improved writing quality — mark_k · 2026-09-04