Inception CEO Stefano Ermon argues diffusion will beat autoregressive models on inference efficiency
No Priors · youtube · 2026-09-18
On the No Priors podcast, Stanford professor and diffusion pioneer Stefano Ermon, co-founder and CEO of Inception, explains why diffusion will win AI inference.
Key points
- Autoregressive LLMs hit hardware and latency bottlenecks; diffusion's parallel token generation offers better inference scaling and GPU utilization on standard hardware.
- Inception extends diffusion beyond images and video to discrete text and code generation with its Mercury models, already deployed in real-world voice agent applications.
- Discussion covers the software stack for serving diffusion models at scale, data compression and structure, controllability, and emergent capabilities.
- Ermon argues the next era of AI competition will be defined by efficiency, touching on future workload splits and academia's role at the frontier.
More from Infra
- Starlink connects thousands of rural Latin American schools, from 10,000 antennas in Honduras to Bolivia's national rollout — XFreeze · 2026-09-18
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18
- TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min — _lewtun · 2026-09-18
- Payments firms race to own AI inference: Stripe taps OpenRouter, Ramp enters the chain — xkonjin · 2026-09-18
- OpenDCAI/DataFlow: open-source pipeline toolkit for pre-training data prep — Puzzleheaded_Box2842 · 2026-09-18
- Qdrant wraps 4+ hour Vector Space Stream on vector search — recording now live — qdrant_engine · 2026-09-18