Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models
TencentARC · hf · 2026-09-22
WorldCrafter from TencentARC is a video world model with a camera-queryable implicit 3D-aware memory for long-horizon, multi-viewpoint consistency. A memory encoder and pose-conditioned readout, trained jointly with the video generator, compress historical observations into a fixed set of target-view tokens before denoising—no explicit depth correspondences. Combined with temporal context and few-step distillation, it enables streaming scene exploration from a single image or text prompt, with substantial gains in long-horizon consistency and camera-control accuracy over minute-scale exploration.
More from Multimodal
- Pixel-art tour of London generated with MiniMax H3 shows consistent stylized video — Sufficient_Cause_43 · 2026-09-22
- Hunyuan Image 3.5 lands exclusively on OnSolo with 5 refs, 2K output, 1 credit — LearnWithBishal · 2026-09-22
- Reddit Asks for Real-World LongCat-Video Inference Times on RTX 4090 to H100 — Clean_Extreme_3970 · 2026-09-22
- BUPT study: RoPE attention decay causes video diffusion models to violate physics — BUPT-CIST · 2026-09-22
- Grok 4.7 made this in Blender — demo shows the model driving 3D software — iamfakhrealam · 2026-09-22
- Kyutai releases Voice of Reason, a speech-native reasoning model hitting 77.1% on GSM8K — alexcovo_eth · 2026-09-22