VGGRPO aligns video diffusion models in latent space with geometry rewards
量子位 · wechat · 2026-07-29
VGGRPO uses latent-space geometry rewards to improve video generation consistency
Researchers from Google, the University of Copenhagen, Oxford and others propose VGGRPO (Visual Geometry GRPO), a post-training framework for video diffusion models that aims to reduce geometric drift, unstable camera motion, and scene-structure inconsistency.
What it does
- Introduces a Latent Geometry Model (LGM) that connects VAE latents to a pretrained geometry foundation model, so geometry can be predicted directly in latent space.
- Replaces expensive RGB-space reward computation with latent-space GRPO.
- Uses two rewards:
- Camera Motion Smoothness Reward to penalize shaky trajectories.
- Geometry Reprojection Consistency Reward to enforce cross-view scene coherence.
Why it matters
The method is designed to preserve the generative strength of pretrained models while improving world consistency, especially in dynamic scenes. The authors argue this is useful for world models and embodied AI, where stable scene understanding is important for prediction, planning and interaction.
The work is listed as accepted at ECCV 2026 and the paper link is provided in the post.
Related event: Google and Partners Introduce VGGRPO to Improve Video Generation(2 posts)→
More from Multimodal
- Video Model WAN 3.0 Now Available on Magnific — aziz4ai · 2026-08-24
- H3 Ref2V Video Gen Test: Footage appears too dark on 4090 — Jeffu · 2026-08-24
- AI generates a Samurai with Bruce Lee's face, highlighting logic failure — Rateko_II · 2026-08-24
- McByte sets SOTA on SportsMOT using segmentation masks for tracking — huggingface · 2026-08-24
- Creator Shares Claude-Generated Comics and Website Update — voooooogel · 2026-08-24
- User generates Red Alert-style RTS assets and animations using Kimi K3 — tinyfool · 2026-08-24