Princeton's VideoGen-Agent Adds 19 Points via Agentic RL Tool Use for Video

princetonu · hf · 2026-09-22

Video generative models still struggle with prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. Princeton's VideoGen-Agent is a multimodal agent trained via multitask agentic RL to orchestrate augmentation, generation, and verification tools through multi-turn interactions. Training combines SFT on teacher trajectories with RL using a category-aware hybrid reward covering tool-call validity, task-appropriate tool use, and video quality.

They also introduce VABench, a 600-prompt held-out benchmark spanning procedural knowledge, single/multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. Results: +19.1 points over the base text-to-video model (56.5→75.6); upgrading generation tools without retraining reaches 86.1; human raters prefer the upgraded configuration in 84.3% of comparisons — showing the trained tool-use policy benefits for free from better generation models.

Original post →

More from Multimodal

Multimodal channel →