4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Yudong Jin, Tao Xie, Qihang Zhang, Zehong Shen, Zhen Xu, Yujun Shen, Hujun Bao, Xiaowei Zhou, Yinghao Xu
cs.CV
2026-08-21
From one uncalibrated clip, 4DAnyone synthesizes 16 consistent views and lifts them into 4DGS. DNA-Rendering reconstruction PSNR is 24.15 versus 20.55 for fine-tuned ReCamMaster.
4D Gaussian Splatting can render a moving person from any viewpoint, but building the representation still wants a calibrated camera array. A phone video has unknown intrinsics and poses. Fitting 4DGS to that clip leaves the back of the body and occluded regions badly under-constrained.
Camera-controlled video diffusion looks like a workaround: synthesize novel views, then reconstruct. Those models look fine at a handful of views. At the tens of views 4DGS actually needs, appearance drifts and structure splits across batches. The DiT attention window is finite, so target views must be denoised in groups. Feeding every previously generated view as reference grows as O(N) and dilutes appearance cues. Groups that never attend to each other slowly disagree on global layout.
4DAnyone sits on Wan2.2-TI2V-5B and follows generate-then-reconstruct. Geometry is sparse on purpose. GVHMR estimates a 3D skeleton from the monocular clip; that skeleton is rasterized at each target camera with a z-buffer, so nearer limbs occlude farther ones. A 2D stick figure cannot tell an arm in front of the torso from one behind it. Only 40 keypoints are kept (body, feet, palm-level hands). Face and finger joints are left to the source video, because noisy fine keypoints hurt more than they help.
Two modules keep many views consistent.
Training has three stages: skeleton-conditioned camera control on DNA-Rendering foregrounds, then unmasked multi-view data for lighting, then monocular Pexels and TedTalk clips for in-the-wild behavior. MVGameHuman, captured in an in-house game engine with 24 cameras, supplies extra diversity. The loss is latent flow matching plus LPIPS at weight 0.25. Training ran on 128 H20 GPUs for about three days. FreeTimeGS lifts the generated videos into 4DGS.
Ten held-out DNA-Rendering scenes and three out-of-distribution DyMVHumans scenes. One frontal source video, sixteen uniformly spaced targets.
| Setting | Method | DNA PSNR | DyMV PSNR |
| Generated-video consistency | ReCamMaster (same-data fine-tune) | 21.47 | 21.94 |
| Generated-video consistency | 4DAnyone | 24.33 | 24.48 |
| 4DGS reconstruction | ReCamMaster | 20.55 | 19.86 |
| 4DGS reconstruction | 4DAnyone | 24.15 | 23.28 |
Zero-shot TrajectoryCrafter lands at 14.81 PSNR on DNA reconstruction and often collapses on front-to-back turns. MV-Performer is plausible from the front and warped on the sides. Ablation without RCP and TCR drops consistency PSNR to 21.09; the full sliding variant reaches 22.63. Random routing adds nothing. Strided routing hurts. Local adjacency during grouping is load-bearing.
This is a usable monocular-to-4D human path that does not estimate dense metric depth or source-camera poses. Skeletons are stabler than wild depth. RCP and TCR treat the bounded attention window as an engineering constraint, not a dead end. Code and a project page are public. For volumetric video, this is generation-assisted reconstruction, not another camera ControlNet stacked on a U-Net.
The bill is real: 128 GPUs, three days, 20 denoising steps, four views per group. ReCamMaster trained with the same RCP, TCR, and data still trails, so the skeleton condition and packing are doing work beyond the dataset.
Flowing fabric far from the body is outside the skeleton's support and becomes inconsistent across views. When HMR misses an unusual pose, every generated view copies the error: an en-pointe dancer is flattened to flat feet all the way around. The test set is small (10 DNA scenes, 3 DyMV). In-the-wild evidence is mostly qualitative. TrajectoryCrafter and MV-Performer are zero-shot, while ReCamMaster is the matched fine-tune; the paper marks that distinction. Synthetic humans can be misused. The authors ask that outputs be labeled synthetic and made only with consent.