VGGRPO Stabilizes Video Geometry via Latent Space RL

机器之心 · wechat · 2026-07-17

Core Issue

While large-scale video diffusion models have significantly improved visual quality, they frequently suffer from geometry drift, camera jitter, and inconsistent scene structures. This makes them particularly unsuitable for downstream tasks like embodied AI and world models that require stable 3D consistency.

What VGGRPO Does

Researchers introduced VGGRPO (Visual Geometry GRPO), which focuses on geometry-aware video post-training directly within the latent space, rather than decoding to RGB first to calculate rewards. It consists of two components:

Results and Significance

Experiments show that VGGRPO outperforms baselines in geometric metrics across both static and dynamic scenes without sacrificing general video quality. The authors emphasize that this method requires no complex changes to the generator architecture and doesn't rely on expensive pixel-level rewards to enhance the model's ability to build a "physically consistent world." This offers significant value for world models, embodied AI, and video generation.

Original post →

More from Multimodal

Multimodal channel →