IFM AI's Uno couples diffusion with LLMs for lossless 2.2x faster generation
srchvrs · x · 2026-09-18
IFM AI introduced Uno, a diffusion-augmented LLM claiming autoregressive quality at diffusion speed, targeting the token-by-token inference bottleneck. It's pitched as a lossless speedup that doesn't degrade response quality.
On K2-Horizon-7B, Uno reportedly beats state-of-the-art diffusion methods on both quality and throughput, with up to 2.2x speedup at no quality cost. Paper and model weights are public.
The reposter adds that Uno's plug-and-play design could extend to post-training and other complex scenarios, and argues diffusion scaling along both depth (denoising steps) and length (parallel test-time scaling) may unlock stronger reasoning.
Related event: IFM AI's Uno: Diffusion-Augmented LLM Delivers Lossless 2.2x Speedup(2 posts)→
More from Models
- Zero-shot classifiers rebranded as 'decision models' surprises HF engineer — mervenoyann · 2026-10-02
- Developer Says Anthropic's Claudebot Hits His Site 280K+ Times Per Day — TejasKumar_ · 2026-10-02
- Dev releases low-bit Qwen3.8-Flash quant keeping 95% bf16 accuracy at long contexts — Crampappydime · 2026-10-02
- "Adoption Is the Real Model Eval": Benchmarks Mean Nothing If Workers Won't Use It — diegoposts · 2026-10-02
- Qwen3.8-Flash-Next on a single R9700: 863 t/s prefill and 35 t/s decode at 230k context — Designer_Elephant227 · 2026-10-02
- AI Now Beats Licensed CPAs on Speed and Accuracy, but Still Can't Close the Books — The Decoder · 2026-10-02