Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models

TencentARC · hf · 2026-09-22

WorldCrafter from TencentARC is a video world model with a camera-queryable implicit 3D-aware memory for long-horizon, multi-viewpoint consistency. A memory encoder and pose-conditioned readout, trained jointly with the video generator, compress historical observations into a fixed set of target-view tokens before denoising—no explicit depth correspondences. Combined with temporal context and few-step distillation, it enables streaming scene exploration from a single image or text prompt, with substantial gains in long-horizon consistency and camera-control accuracy over minute-scale exploration.

Original post →

More from Multimodal

Multimodal channel →