Tencent's WorldCycle: Self-Verifiable RL for Long-Horizon Video World Models

tencent · hf · 2026-08-06

To address compounding errors in long-horizon planning within interactive video world models, Tencent introduced WorldCycle, a self-verifiable reinforcement learning framework.

The key insight is that reversible action cycles enable verification: a sequence combined with its inverse must analytically return to the initial state, providing annotation-free supervision on long-horizon correctness. WorldCycle optimizes two complementary rewards:

This forces the model to learn actions as consistent state operators rather than memorizing temporal patterns. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x. The team also released CycleBench, a diagnostic benchmark for state-returning abilities.

Original post →

More from Research

Research channel →