Uno hybrid diffusion LLM claims 'beats all', but latency-quality is Pareto dominated
joao_gante · x · 2026-09-04
Researcher ssahoo introduced "Uno", targeting diffusion LLMs' two weaknesses vs AR models: lower quality and slower large-batch inference.
- Keeps the AR architecture, with each layer holding two weight sets: AR weights and Diffusion weights
- Diffusion weights enable lossless parallel sampling from the AR distribution
- Claims: faster than all speculative decoding methods (DFlash, EAGLE-3); beats all diffusion LLMs (Mercury 2, Diffusion Gemma, Llada); paper, models, and code released
Contested: bodonoghue85 says the "beats ALL" claim is false — Uno-Qwen is comfortably Pareto dominated on latency-quality by both Mercury 2 and DiffusionGemma, and Uno's LCBv6 number is oddly missing from the writeup.
Related event: Uno: diffusion-augmented LLM claims AR-level quality with faster inference(3 posts)→
More from Research
- DeepMind's "LLM can't jump" argument: induction and deduction can't produce scientific revolutions — 0xsachi · 2026-09-04
- Free 39-Episode Control Bootcamp: The Control Theory That Runs Real Robots — lukas_m_ziegler · 2026-09-04
- UniReps Workshop in Paris opens call for papers on unified representations, due Oct 4 — ClementineDomi6 · 2026-09-04
- Utopia: open-source bitemporal knowledge graph gives RAG agents a memory of change — Shruti_0810 · 2026-09-04
- Emotion is an optimizer's control plane, not an irrational advisor — mimi10v3 · 2026-09-04
- New f-loss Cures Spectral Bias in Pixel-Space Flow Matching, Speeding Convergence — serrjoa · 2026-09-04