Uno hybrid diffusion LLM claims 'beats all', but latency-quality is Pareto dominated

joao_gante · x · 2026-09-04

Researcher ssahoo introduced "Uno", targeting diffusion LLMs' two weaknesses vs AR models: lower quality and slower large-batch inference.

Contested: bodonoghue85 says the "beats ALL" claim is false — Uno-Qwen is comfortably Pareto dominated on latency-quality by both Mercury 2 and DiffusionGemma, and Uno's LCBv6 number is oddly missing from the writeup.

Related event: Uno: diffusion-augmented LLM claims AR-level quality with faster inference(3 posts)→

Original post →

More from Research

Research channel →