Caltech's DynaTokens teaches dynamics to camera-controlled video models at test time, NeurIPS accepted
georgiagkioxari · x · 2026-09-30
Caltech researchers Ziqi Ma, Hongqiao Chen and Georgia Gkioxari introduce DynaTokens, a NeurIPS-accepted fix for camera-controlled video models that nail camera motion but fail at scene dynamics:
- Key insight: camera control is global (affects every patch in every frame) while scene dynamics are local ("a dog running" should only affect the dog); global adaptation methods like LoRA, block finetuning and TTT entangle the two, so dynamics supervision degrades camera control
- Method: a lightweight add-on that trains a few learnable, scene-specific tokens at test time, decoupling object motion from camera motion
- Results: outperforms HYWP, Lyra-2, Lingbot and dynamics-focused LiveWorld/HyDRA on realistic scenes, and generalizes to unseen camera trajectories (WASD translation + arrow-key rotation).
Project page, paper and code are public.
More from Multimodal
- Opus 5.5 + ElevenLabs Turn a Treasury Figure Into Rock and Rap Tracks — leveragedupside · 2026-09-30
- Lumibelle: open-source video editor feeds reusable character reels into MiniMax H3 for consistency — rad_reverbererations · 2026-09-30
- Asking for help: outpainting 9:16 video to 16:9 with Minimax H3 changes the footage entirely — navarisun · 2026-09-30
- Inception Labs launches Mercury Voice, a diffusion LLM with 2x lower latency than GPT-6 Luna and Claude Haiku 4.5 — StefanoErmon · 2026-09-30
- ComfyUI-Continuity Update Adds Cast Management, RefMod Presets and Seam Optimization — Okims_kor · 2026-09-30
- AI turns its subagents into a rock band, full MV made with local music and video models — WolframRvnwlf · 2026-09-30