Stanford's 8-lecture diffusion course, condensed: one teddy bear prompt through every step

le_james94 · x · 2026-09-22

James Le published a full recap of Stanford CME 296 (Diffusion & Large Vision Models, taught by Afshine and Shervine Amidi), 14 hours across 8 lectures that follow a single teddy-bear prompt from noise to finished image through autoencoders, denoising transformers, guidance, and evaluation. He also argues decomposed metrics only become objective once an AI judge outperforms human-human agreement (VIEScore+GPT-4v: 0.3 Spearman vs 0.45 human-human), and flags an open question: nobody measures how much of a generated image comes from the generator vs the VAE decoder.

Original post →

More from Research

Research channel →