GUSH3R: Photorealistic Dynamic Human Scene Reconstruction from Monocular Video

janusch_patas · x · 2026-07-07

The paper introduces GUSH3R, tackling the novel problem of feed-forward, photorealistic, and renderable dynamic human scene reconstruction from monocular video. Core contributions include an architecture bridging human scene foundation models with photorealistic rendering, utilizing geometric priors and SMPL-X representations. It achieves competitive novel view synthesis quality compared to decomposition-based and optimization-based baselines while significantly boosting inference efficiency.

Related event: GUSH3R Enables Dynamic 3D Reconstruction from Monocular Video(2 posts)→

Original post →

More from Multimodal

Multimodal channel →