HKU and Tencent Introduce SCoPE for 3D-Aware Video World Models

jiqizhixin · x · 2026-08-29

HKU, HKUST, and Tencent ARC Lab present SCoPE (Sightline-Coordinate Positional Encoding), improving positional encoding in video DiTs.

Core Improvement:

Results: Significantly improves camera control, cross-view consistency, and revisit consistency. As model scale grows from 5B to 14B, SCoPE's lead widens, with translation error advantage growing from 12% to 25% and FVD improvement from 23% to 47%.

Original post →

More from Research

Research channel →