Tencent ARC open-sources WorldCrafter, a video world model with implicit 3D-aware memory

yshan2u · x · 2026-09-23

Tencent ARC's WorldCrafter skips explicit 3D reconstruction entirely: the video world model learns an implicit, camera-queryable 3D-aware memory, so pointing the camera anywhere lets the memory fill in the view.

Key insight: the requested viewpoint shapes how multi-view evidence gets compressed into the video generator's limited token budget — a jointly trained memory encoder and pose-conditioned readout produce a fixed set of target-view tokens before denoising, with no explicit depth correspondence. Combined with recent temporal context and few-step distillation, it enables streaming exploration of a scene from a single image or text prompt, with minute-scale horizon consistency and strong camera control.

Weights (base + distilled fast version) are on Hugging Face, code on GitHub, paper at arXiv:2609.24984. Resolution is 384×640 for now, expected to scale.

Related event: TencentARC Open-Sources WorldCrafter, a Video World Model with Implicit 3D-Aware Memory(9 posts)→

Original post →

More from Multimodal

Multimodal channel →