Stanford's 8-lecture diffusion course, condensed: one teddy bear prompt through every step
le_james94 · x · 2026-09-22
James Le published a full recap of Stanford CME 296 (Diffusion & Large Vision Models, taught by Afshine and Shervine Amidi), 14 hours across 8 lectures that follow a single teddy-bear prompt from noise to finished image through autoencoders, denoising transformers, guidance, and evaluation. He also argues decomposed metrics only become objective once an AI judge outperforms human-human agreement (VIEScore+GPT-4v: 0.3 Spearman vs 0.45 human-human), and flags an open question: nobody measures how much of a generated image comes from the generator vs the VAE decoder.
More from Research
- TinyTorch: PyTorch's free curriculum to build an ML framework from scratch in 20 modules — PyTorch · 2026-09-22
- Why AI won't boost paper output for researchers who chase hard problems — kfountou · 2026-09-22
- ICML 2027 braces for 100k submissions as AI paper boom continues — CharlotteHase · 2026-09-22
- 1080 Ti beats RTX 6000 by 2.4x on dense-model inference despite 4x less bandwidth — EAccelerate_42 · 2026-09-22
- Clem Delangue: open RL environments are the new open pretraining data — Thom_Wolf · 2026-09-22
- Harrison Chase on using decision models to rein in agent workflows — hwchase17 · 2026-09-22