LightInteraction speeds interactive video world models up 2.59× without retraining

新智元 · wechat · 2026-07-22

LightInteraction cuts interactive video generation latency with state-aware inference

Researchers from Zhejiang University and NVIDIA propose LightInteraction, a training-free inference method for interactive video world models. The goal is to make long-horizon, game-like video generation faster and more stable without changing model parameters.

The paper identifies three main bottlenecks in interactive video generation:

LightInteraction turns the camera trajectory into a scheduling signal and adapts computation to the interaction state:

On HY-WorldPlay, the method reduces main-generation latency from 228.60s to 88.24s for about 10 seconds of video, a 2.59× speedup, while peak memory drops from 76.57GB to 54.66GB. On Matrix-Game-3.0, latency falls from 59.70s to 37.07s for about 20 seconds of video.

The key takeaway is that interactive video systems can be accelerated by making inference state-aware, not just by shrinking the model.

Original post →

More from Multimodal

Multimodal channel →