Controlled study finds training data quality is decisive for text-to-video models
Amber Yijia Zheng · hf · 2026-07-23
Moving Alphabet studies how training data affects text-to-video generation, using a controlled procedural testbed.
The paper renders letters with controlled variation in font, color, size, position, direction, and speed, then corrupts metadata to isolate the impact of data distribution and caption quality. The main findings are:
- Balanced, diverse video content and duration distribution are critical for generalization.
- Caption quality strongly affects both performance and training efficiency, suggesting text-to-video models are limited by video understanding.
- Classifier-free guidance and fine-tuning on cleaner data help recover some performance, but cannot fully fix poor pre-training data.
The authors argue that pre-training data deserves more scientific attention in large-scale video generation.
More from Multimodal
- FLUX 3 can invent its own camera cuts from a one-line image prompt — venturetwins · 2026-07-23
- Black Forest Labs announces FLUX 3, a multimodal model for image, video, and audio — chrisfirst · 2026-07-23
- Self-Flow speeds multimodal model convergence by up to 2.8×, paper says — hila_chefer · 2026-07-23
- Viewer gifts now trigger real-time AI video effects in a live-streaming demo — ming_calligraphy · 2026-07-23
- GLM-5.2 vision model baseten/GLM-5.2-Vision-NVFP4 trends on Hugging Face — baseten · 2026-07-23
- FLUX 3 is announced, but its capabilities will roll out over weeks and months — Angaisb_ · 2026-07-23