Controlled study finds training data quality is decisive for text-to-video models
Amber Yijia Zheng · hf · 2026-07-23
Moving Alphabet studies how training data affects text-to-video generation, using a controlled procedural testbed.
The paper renders letters with controlled variation in font, color, size, position, direction, and speed, then corrupts metadata to isolate the impact of data distribution and caption quality. The main findings are:
- Balanced, diverse video content and duration distribution are critical for generalization.
- Caption quality strongly affects both performance and training efficiency, suggesting text-to-video models are limited by video understanding.
- Classifier-free guidance and fine-tuning on cleaner data help recover some performance, but cannot fully fix poor pre-training data.
The authors argue that pre-training data deserves more scientific attention in large-scale video generation.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11