4DAnyone: Reconstruct a 4D human from a casual phone video, no rig needed
anselm · x · 2026-08-31
4DAnyone (Zhejiang University, Ant Group et al., SIGGRAPH Asia 2026) reconstructs a 4D Gaussian Splatting volumetric human from a single casually captured monocular video — no rig, no calibration, no tripod.
- Traditional photorealistic 4D human capture requires a calibrated synchronized camera array; this work instead generates the videos such an array would have recorded.
- Key challenge is consistency at scale: Reference Context Packing (RCP) compresses linearly growing reference context into a fixed budget, while Target Context Routing (TCR) routes context across disjoint generation groups to prevent structural drift; a 3D skeleton supplies sparse-but-accurate conditioning.
- Generalizes robustly to dance, sports, stage, speech and fashion footage, tolerating mild camera motion and unknown intrinsics/poses.
- Paper, code and models are released.
Related event: 4DAnyone: Reconstructing 4D Humans from a Single Phone Video(3 posts)→
More from Multimodal
- Breeze-TTS-2 demo now available on Hugging Face — BreezeBlue · 2026-09-01
- 求助:Anima 模型训练正常但推理生成模糊 — a_throwawayorsmthn · 2026-09-01
- 求测:Ideogram 4 INT8 量化版在 3060 12G 上的表现 — WhyDoiHearBosssMusic · 2026-09-01
- Help: Swapping a Mustache Using a Reference Image in Flux Workflows — sadboi2021 · 2026-09-01
- Fal's new model heralds next chapter for Hollywood VFX — briannekimmel · 2026-09-01
- Sea Angels Generated in p5.js Using Claude Opus 5 — anselm · 2026-09-01