Discrete diffusion delivers provably lossless LLM inference speedups, drop-in for training
Cohere · youtube · 2026-10-03
A talk in the Cohere Labs in Conversation series, presented by Subham Sahoo, Sr. Research Scientist at MBZUAI-IFM, on unlocking lossless speedups in LLMs via discrete diffusion.
- Autoregressive models generate high-quality text but remain slow due to sequential decoding; diffusion language models promise much faster parallel inference, yet current systems like Mercury2, Diffusion Gemma, and Nemotron Diffusion still lag frontier AR models in quality.
- The talk presents a discrete diffusion approach that achieves provably lossless inference speedups: it works as a drop-in replacement for existing training pipelines, preserves model quality, and significantly improves throughput across levels of inference parallelism.
- Experiments show it qualitatively outperforms existing diffusion baselines while running substantially faster than autoregressive decoding.
More from Infra
- Qualcomm's Snapdragon 8 Elite Gen 6 hits 5GHz, runs 30B+ param MoE models on-device — DavidLinthicum · 2026-10-03
- DFlash2 speculative decoding hits 2.8x speedup in local Qwen3.8 three-way benchmark — FantasticNature7590 · 2026-10-03
- Cactus releases Whistle: a 16.9MB speech-to-text model that beats Whisper base on CPU with 6x speed — ycombinator · 2026-10-03
- Cohere and vLLM co-host Toronto meetup on open weights and inference — cohere · 2026-10-03
- TSMC evaluating multi-billion dollar Texas fab campus amid surging US demand from Nvidia, Apple — Beth_Kindig · 2026-10-03
- Runner pays $4,500 for a third GPU to run near-frontier models locally — marian_nmt · 2026-10-03