Controlled study finds training data quality is decisive for text-to-video models
Amber Yijia Zheng · hf · 2026-07-23
Moving Alphabet studies how training data affects text-to-video generation, using a controlled procedural testbed.
The paper renders letters with controlled variation in font, color, size, position, direction, and speed, then corrupts metadata to isolate the impact of data distribution and caption quality. The main findings are:
- Balanced, diverse video content and duration distribution are critical for generalization.
- Caption quality strongly affects both performance and training efficiency, suggesting text-to-video models are limited by video understanding.
- Classifier-free guidance and fine-tuning on cleaner data help recover some performance, but cannot fully fix poor pre-training data.
The authors argue that pre-training data deserves more scientific attention in large-scale video generation.
More from Multimodal
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11