Video diffusion still leans on 8- to 11-year-old evaluation metrics
kalomaze · x · 2026-07-23
The post argues that much of the video-diffusion world-modeling discussion is being judged with stale metrics and stale feature extractors.
The image highlights how old the key evaluation pieces are: Inception-v3 dates to 2015, FID to 2017, I3D to 2017, and FVD to 2018. The point is that many baselines are still evaluated using classifiers and metrics that were designed years ago, raising questions about whether current progress is being measured with outdated tooling.
Related event: Diffusion Model Research Still Relies on Outdated Metrics(2 posts)→
More from Research
- AgenC now defaults to one agent after multi-agent systems fell 39%–70% behind — tetsuoai · 2026-07-23
- AgenC defaults to one agent after papers found multi-agent swarms underperformed — tetsuoai · 2026-07-23
- FLOC26 panel will discuss what AI progress means for CS and formal methods — swarat · 2026-07-23
- LLMs may help most in the manual refinement phase of decompilation — OwariDa · 2026-07-23
- RLSS 2026 in Milan turns Bellman backups into a coffee-fueled group sport — misovalko · 2026-07-23
- RLSS 2026 Milan masterclass on world models and RL spans 23 slides — misovalko · 2026-07-23