Inception CEO Stefano Ermon bets on diffusion LLMs that generate tokens in parallel
saranormous · x · 2026-09-21
The latest No Priors episode features Stefano Ermon, CEO of Inception AI, on why true architectural AI research is rare today and why it's still worth pursuing good ideas in the age of scaling.
- Core thesis: today's LLMs can't generate token 10 until token 9 exists due to autoregression; Ermon bets on diffusion LLMs that generate tokens in parallel.
- The conversation also covers his company's latency focus and chasing fundamentally new architectures while scaling dominates.
- A direct take on whether architectural innovation still has room beyond scaling.
Related event: Stefano Ermon Bets on Diffusion LLMs to Replace Token-by-Token Decoding(2 posts)→
More from Models
- YOCO back in spotlight: blog breaks down cross-layer KV sharing in DeepSeek-V4.1-Flash and Gemma 4 — donglixp · 2026-09-21
- Gemini 4.0 Rumored Next: Yearly Pro Updates, Monthly Flash Releases Expected — haider1 · 2026-09-21
- SemiAnalysis' Dylan Patel: Chinese labs going closed-source, 'open is dying quickly' — ns123abc · 2026-09-21
- DeepSeek Web Output Style Reportedly Lobotomized by Safe Harbor RLHF — Old_Let6328 · 2026-09-21
- JEV opens to all with $5 free credits; $0.042 input and free output pricing — op7418 · 2026-09-21
- ChatGPT 20x Max users hit tighter limits, suspect compute favors government and API customers — BopSupreme · 2026-09-21