JHU's GenCine turns a single image into an editable 3D scene for camera and object motion control in video generation
anand_bhattad · x · 2026-10-03
- Johns Hopkins researchers (led by Jiahan Zhang, with Alan Yuille and Anand Bhattad) present Generative Cinematographer (GenCine), arXiv:2610.02180.
- Problem: existing controllable video generation relies on 2D trajectories or drag signals, which are ambiguous when camera and objects move simultaneously.
- Method: lift a single image into an editable 3D point cloud, author camera paths and foreground motion via local 3D handles in tools like Blender, then project controls into correspondence maps (world-space XYZ + identity) read by a frozen Wan VAE, injected into a pretrained video diffusion model via a lightweight side branch and LoRA adapters.
- Authors say the 3D composition could extend beyond cinematography to robotics training data and viewpoint exploration.
Related event: GenCine: Single-Image Video Generation with 3D Motion Control(3 posts)→
More from Multimodal
- Creator shares a favorite StyleGAN model: "such a trip" — makeitrad1 · 2026-10-03
- HeyGen open-sources HyperFrames: deterministic HTML-to-MP4 video pipelines with an agent skill pack — lmoroney · 2026-10-03
- fal launches MiniMax H3 Max Recast: swap video subjects from a reference photo at $0.30/s, motion and audio intact — lmoroney · 2026-10-03
- Qwen-Image-2.1-viggle-turbo: fast 6-step character sheets in ComfyUI — RiverSide71h · 2026-10-03
- Camera Control for AI Images in ComfyUI with QI 2.1 — RobbaW · 2026-10-03
- GenCine comparisons show 3D controls guiding a robot arm where baselines barely move — anand_bhattad · 2026-10-03