Controlled study finds training data quality is decisive for text-to-video models

Amber Yijia Zheng · hf · 2026-07-23

Moving Alphabet studies how training data affects text-to-video generation, using a controlled procedural testbed.

The paper renders letters with controlled variation in font, color, size, position, direction, and speed, then corrupts metadata to isolate the impact of data distribution and caption quality. The main findings are:

The authors argue that pre-training data deserves more scientific attention in large-scale video generation.

Original post →

More from Multimodal

Multimodal channel →