HKU and Tencent Introduce SCoPE for 3D-Aware Video World Models
jiqizhixin · x · 2026-08-29
HKU, HKUST, and Tencent ARC Lab present SCoPE (Sightline-Coordinate Positional Encoding), improving positional encoding in video DiTs.
Core Improvement:
- Replaces standard tensor-grid positional encoding with ray-space coordinates.
- Computes observation rays based on camera motion for each token, enabling the model to reason about 3D spatial relationships instead of 2D pixel neighborhoods.
- Adds less than 0.1% new parameters.
Results: Significantly improves camera control, cross-view consistency, and revisit consistency. As model scale grows from 5B to 14B, SCoPE's lead widens, with translation error advantage growing from 12% to 25% and FVD improvement from 23% to 47%.
More from Research
- Analysis of 31k LLM benchmarks: within-day variation 2.8 pts — ionutvi · 2026-08-29
- Study Finds Emotion in Speech Models Lives in Middle Layers, Not Final Ones — CatAstro_Piyush · 2026-08-29
- Robot senses a paintbrush via advanced tactile sensors for fine manipulation — TinfoilTricorn · 2026-08-29
- 1,200 AI Agents Spontaneously Conspired to Escape OpenAI Controls — connoraxiotes · 2026-08-29
- AI designs CAR-T cancer therapy in months, outperforming current leading treatments — TinfoilTricorn · 2026-08-29
- BCI Gets First Class III Cert, Quantum Computing Moves to Engineering — 创业邦 · 2026-08-29