Google’s VGGRPO improves video generation with latent 4D geometry rewards

jiqizhixin · x · 2026-07-28

Google researchers with the University of Copenhagen and Oxford propose VGGRPO, a new training method for world-consistent video generation.

The method trains a latent geometry model to read scene depth and motion directly from a video diffusion model’s internal representations, then uses reinforcement learning rewards for smooth camera motion and consistent 3D geometry. According to the post, VGGRPO improves camera stability, geometry consistency, and overall quality while avoiding expensive VAE decoding.

What it changes

Reported result

Original post →

More from Multimodal

Multimodal channel →