ShadowDancer: Teaching Video World Models Any Action via Shadow Pairs
AlayaLab · hf · 2026-07-31
ShadowDancer introduces a novel approach for frame-level control of interactive video world models, enabling any-action generation.
- The Challenge: Existing interfaces either encode actions too loosely (letting models improvise) or rely on hard-to-acquire structured signals, making precise control across diverse dynamics impractical.
- Key Innovations:
- Shadow Pairs: Video pairs that replay the same dynamics under independently resampled appearances, constructed at scale to enable exact control.
- Cross-Shadow Prediction: Learns actions by predicting one shadow from another, discarding resampled variables to extract a unified dynamics representation.
- Results: Turns any demonstrated clip into a reusable action asset without needing action labels, motion estimators, or fine-tuning. Achieves an 86% average blinded win rate over strong baselines.
More from Multimodal
- Google open-sources Glanceboard: Turn your calendar into e-ink art with Gemini — osanseviero · 2026-07-31
- Open Source Glanceboard: Turn Calendars into E-ink Art with Gemini — osanseviero · 2026-07-31
- Video Relighting Approaches Compared: LTX-LoRA Renders in 2 Minutes — mickmumpitz · 2026-07-31
- MiniMax H3 video model hits Topview at just 30% of Seedance's price — iamfakhrealam · 2026-07-31
- ComfyUI-LTX-Reframe Nodes Released: Canvas & Audio Workflow — Championship_Better · 2026-07-31
- Generating 5s Video with LTX2 on RTX 3060 Takes 10 Mins: How to Optimize? — Nicola883 · 2026-07-31