Check per-timestep loss curves in small diffusion training runs
jfischoff · x · 2026-09-06
A reminder for small-scale diffusion training: look at per-timestep loss curves rather than only the aggregate loss, to catch issues in specific denoising stages.
More from Research
- DeepMind paper shows cheating spreading like an epidemic across ~100 AI agents — jackclarkSF · 2026-09-06
- Context Compaction Theory: first formal proof linking agent compaction to communication complexity — lateinteraction · 2026-09-06
- Last theorem on Freek Wiedijk's famous list has been formalized — satnam6502 · 2026-09-06
- New theory shows how to scale residual network updates when layer weights are correlated — burkov · 2026-09-06
- Terence Tao post sparks buzz as a cool mechanistic interpretability application — Sauers_ · 2026-09-06
- GPT-6 Astra tops surgical AI leaderboard but loses to models 1000x smaller — ddonoho · 2026-09-06