Google’s VGGRPO improves video generation with latent 4D geometry rewards
jiqizhixin · x · 2026-07-28
Google researchers with the University of Copenhagen and Oxford propose VGGRPO, a new training method for world-consistent video generation.
The method trains a latent geometry model to read scene depth and motion directly from a video diffusion model’s internal representations, then uses reinforcement learning rewards for smooth camera motion and consistent 3D geometry. According to the post, VGGRPO improves camera stability, geometry consistency, and overall quality while avoiding expensive VAE decoding.
What it changes
- Adds a latent geometry model on top of video diffusion latents
- Uses two rewards: motion smoothness and geometry consistency
- Works entirely in latent space
Reported result
- Better camera stability
- Better geometry consistency
- Better overall generation quality
More from Multimodal
- Google DeepMind reconstructs Pelé’s lost 1959 goal and opens GNM to developers — plopesresearch · 2026-07-28
- Krea 2 reportedly generated all 50 car models from a single prompt set — poopoo_fingers · 2026-07-28
- Opus 5 demo shows product videos with smooth camera moves and zero editing — iamrobotbear · 2026-07-28
- Diffusers adds Nunchaku 4-bit diffusion inference with up to 50% less VRAM — dl_weekly · 2026-07-28
- Kimi-K3 adds image-text-to-text support at 8-bit precision on GPU stacks — petrusenko_max · 2026-07-28
- Showcase: "A Massa" AI-Generated Animated Music Video Inspired by Brazilian Culture — LeoMendesMusic · 2026-07-28