A New Method to Make Video Models Obey Better
Cohere · youtube · 2026-07-18
This presentation discusses how to make video generation models better follow user intent, focusing on improving obedience to control boxes via minimal adjustments.
Core Method
- Instead of retraining an entire text-to-video diffusion model, it optimizes the user-provided control boxes
- Aligns control boxes with the model's internal attention mechanism
- Uses smooth masking to stabilize the optimization process
Results
- Achieves better control compared to existing baselines on the AnimalKingdom dataset
- Takeaway: even minor adjustments can significantly improve spatial constraint adherence in video generation
Speaker Background
Daniel Ajisafe is a CS PhD student at the University of British Columbia researching controllable generative models and user intent alignment. He is an RBC Borealis AI Fellow and a Best Paper award recipient.
More from Multimodal
- Reddit shares an AI-generated mini movie called The Lunar Ship — Ermajean12 · 2026-07-21
- AI creator GossipGoblin is turning short-form clips into a feature film — Hackedv12 · 2026-07-21
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21
- SVG Generation Comparison: Leading AI Models Draw a Red Ferrari — Able-Line2683 · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21