Google’s VGGRPO improves video generation with latent 4D geometry rewards
jiqizhixin · x · 2026-07-28
Google researchers with the University of Copenhagen and Oxford propose VGGRPO, a new training method for world-consistent video generation.
The method trains a latent geometry model to read scene depth and motion directly from a video diffusion model’s internal representations, then uses reinforcement learning rewards for smooth camera motion and consistent 3D geometry. According to the post, VGGRPO improves camera stability, geometry consistency, and overall quality while avoiding expensive VAE decoding.
What it changes
- Adds a latent geometry model on top of video diffusion latents
- Uses two rewards: motion smoothness and geometry consistency
- Works entirely in latent space
Reported result
- Better camera stability
- Better geometry consistency
- Better overall generation quality
Related event: Google and Partners Introduce VGGRPO to Improve Video Generation(2 posts)→
More from Multimodal
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- ComfyUI trick: aux preprocessor + Qwen transfers poses across characters with one prompt — Acceptable-Work8202 · 2026-09-23
- Same portrait prompt across Midjourney V6.1, V7 and V8.2: do older models look better? — tisch_eins · 2026-09-23
- Testing AI character consistency across a 20-image travel sequence — SiennaVaire · 2026-09-23
- Midjourney v8.2 Faces: New Portrait Generation Samples Shared — azed_ai · 2026-09-23
- One-sentence prompt generates lifelike dog video, shown side-by-side with the real one — wgrathwohl · 2026-09-23