Code World Model: LLM agents write the rules, video models render the visuals
新智元 · wechat · 2026-09-07
- Researchers from Westlake University's AGI Lab and NTU propose Code World Model, a new world-model paradigm that splits responsibilities: a coding agent maintains an executable world state via code (goals, rules, long-term causality), while a video model handles fine-grained visual rendering.
- A lightweight "proxy" interface — coarse geometric skeletons at roughly 1/16 the pixel count of target video — bridges the two, providing per-frame spatial constraints.
- Training uses 5.6 hours of gameplay recordings, yielding 9,420 strictly aligned 5-second clips. The prototype fine-tunes a MiniMax-H3 video backbone with rank-128 LoRA (596M trainable params, 8x H800), with GPT-5.6 Sol as the coding agent at inference.
- Despite minimal data, the same proxy skeleton can be rendered as characters of different identities and art styles while obeying specified positions, trajectories, and camera motion. Paper, project page, and code are public.
More from Research
- RoboFollow Benchmark Exposes Instruction-Following Mirage in 9 VLA Policies — AutoLab-SJTU · 2026-09-23
- Cohere Labs session probes scaling law reliability and offers a research checklist — Cohere_Labs · 2026-09-23
- NVIDIA's SoL-Pi lets AI rewrite agent harnesses, cutting tokens 44.7-49% — alex_verem · 2026-09-23
- Slingshot RL framework jailbreaks Qwen2.5-32B at 67% success, transfers zero-shot to Gemini 2.5 Flash — j_foerst · 2026-09-23
- TLAPS-Bench launches in alpha to test whether AI agents can formally prove system correctness — tianyin_xu · 2026-09-23
- ImIR replaces text prompts with image instructions for all-in-one restoration — Süleyman Aslan · 2026-09-23