World Labs' LoGo uses local-global rewards to fix 3D inconsistency in long-horizon camera-controlled video

gowthami_s · x · 2026-10-06

Researchers from World Labs and Caltech (including Li Fei-Fei and Ben Mildenhall) present LoGo, a post-training method that substantially improves 3D consistency in long-horizon, camera-controlled video generation.

Problem: as camera-controlled models extend generation horizons, objects lose permanence — scene layouts shift on revisit, objects appear/disappear or deform, and artifacts emerge. Existing post-training assigns a single scalar reward to the whole clip, poorly suited to correcting errors over long horizons.

Method: LoGo blends a spatially localized reward, giving fine-grained credit assignment that markedly improves 3D consistency, with a global reward that preserves camera following and video quality.

Results: across three base models, LoGo clearly outperforms prior methods on DL3DV and TrajectoryBench, a new benchmark for long-horizon complex camera trajectories that current evaluations lack. Paper, code, and benchmark are released.

Related event: World Labs Releases LoGo to Improve 3D Consistency in Long Video Generation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →