Princeton's VideoGen-Agent Adds 19 Points via Agentic RL Tool Use for Video
princetonu · hf · 2026-09-22
Video generative models still struggle with prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. Princeton's VideoGen-Agent is a multimodal agent trained via multitask agentic RL to orchestrate augmentation, generation, and verification tools through multi-turn interactions. Training combines SFT on teacher trajectories with RL using a category-aware hybrid reward covering tool-call validity, task-appropriate tool use, and video quality.
They also introduce VABench, a 600-prompt held-out benchmark spanning procedural knowledge, single/multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. Results: +19.1 points over the base text-to-video model (56.5→75.6); upgrading generation tools without retraining reaches 86.1; human raters prefer the upgraded configuration in 84.3% of comparisons — showing the trained tool-use policy benefits for free from better generation models.
More from Multimodal
- MiniMax H3 Turns a Mac Desktop Into Oggy-Style Cartoon Chaos — Full Prompt Included — SimplyAnnisa · 2026-09-22
- Striking human-motion visualizations made with Sentinel — Kyrannio · 2026-09-22
- New Pika impresses early users as the video startup returns to form — Kyrannio · 2026-09-22
- Tencent's Hunyuan Image 3.5 lands on OnSolo: 5 refs, 2K output, 1 credit — Div_pradeep · 2026-09-22
- Museum statue meme generated with AI — prompt included — umesh_ai · 2026-09-22
- Intel Ships Day-0 OpenVINO Support for Qwen-Image-2.1 — Alibaba_Qwen · 2026-09-22