Tencent ARC open-sources WorldCrafter, a video world model with implicit 3D-aware memory
yshan2u · x · 2026-09-23
Tencent ARC's WorldCrafter skips explicit 3D reconstruction entirely: the video world model learns an implicit, camera-queryable 3D-aware memory, so pointing the camera anywhere lets the memory fill in the view.
Key insight: the requested viewpoint shapes how multi-view evidence gets compressed into the video generator's limited token budget — a jointly trained memory encoder and pose-conditioned readout produce a fixed set of target-view tokens before denoising, with no explicit depth correspondence. Combined with recent temporal context and few-step distillation, it enables streaming exploration of a scene from a single image or text prompt, with minute-scale horizon consistency and strong camera control.
Weights (base + distilled fast version) are on Hugging Face, code on GitHub, paper at arXiv:2609.24984. Resolution is 384×640 for now, expected to scale.
More from Multimodal
- PixVerse world model lets you walk through AI-generated video with WASD — Aiden_Tech_Ai · 2026-09-23
- What do agents think they look like? This project has a painter draw their self-descriptions — Mastbubbles · 2026-09-23
- One person spent months making a solo 20-minute-per-episode AI drama series — Puzzleheaded_Tart824 · 2026-09-23
- Testing Hunyuan Image 3.5 via Miora's Image Generator workflow — HeyAmit_ · 2026-09-23
- Hunyuan Image 3.5 realism test: lighting, materials and camera language — HeyAmit_ · 2026-09-23
- Hands-on: Tencent's Hunyuan Image 3.5 preview impresses with text, references and editing control — HeyAmit_ · 2026-09-23