Inception CEO Stefano Ermon: Diffusion Beats Autoregressive LLMs on Inference
No Priors · rss · 2026-09-18
Stanford professor and diffusion pioneer Stefano Ermon, now CEO of Inception, joins No Priors to argue diffusion will win AI inference. Key points:
- Autoregressive bottleneck: token-by-token generation hits hardware and latency limits, hurting inference scaling.
- Diffusion advantage: parallel token generation delivers superior inference scaling and GPU utilization; Inception extends diffusion beyond images/video to discrete text and code generation.
- Mercury models: details on Inception's diffusion LLMs, real-world voice agent deployments, and the software stack needed to serve diffusion at scale.
- Outlook: the next era of AI competition will be defined by efficiency; also covers academia's frontier role, hiring, and recursive self-improvement.
More from Infra
- Starlink connects thousands of rural Latin American schools, from 10,000 antennas in Honduras to Bolivia's national rollout — XFreeze · 2026-09-18
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18
- TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min — _lewtun · 2026-09-18
- Payments firms race to own AI inference: Stripe taps OpenRouter, Ramp enters the chain — xkonjin · 2026-09-18
- OpenDCAI/DataFlow: open-source pipeline toolkit for pre-training data prep — Puzzleheaded_Box2842 · 2026-09-18
- Qdrant wraps 4+ hour Vector Space Stream on vector search — recording now live — qdrant_engine · 2026-09-18