Code World Model: Coding Agent as the World Brain for Video Generation
Scobleizer · x · 2026-08-28
Westlake University's AGI Lab and NTU propose Code World Model: a language-model-brained paradigm where a coding agent continuously maintains an explicit world state (as code) that then guides a video model to generate high-fidelity visuals.
Two motivations:
- Complex world interactions involve goals, rules, causality and other high-level semantics beyond motion and collisions, requiring LM reasoning.
- Games—the main data source for video world models—are just visual outputs of code execution; learning action-to-video mappings directly forces the video model to implicitly approximate program outputs, which is inefficient and entangles world evolution with visual generation.
A consistent world state unlocks long-horizon generation, demonstrated with periodic style changes every 60 seconds and coherent characters across stylized scenes.
Related event: Code World Model: Coding Agent as the World's Brain(3 posts)→
More from Multimodal
- AI video gen moves to production: 35M downloads, 1.7B views for AI dramas — HeyAmit_ · 2026-08-28
- TontaubeV1 Released: 2.9B Parameter Open TTS for Local Long-form Generation — EAVDR · 2026-08-28
- SpatialCrafter enables consistent video gen from single images via 3D proxies — kwangmoo_yi · 2026-08-28
- Australia bans fully AI-generated songs from official charts — Content-Cheetah-6958 · 2026-08-28
- Hy4 Preview generates detailed Sakura Bonsai image via WorkBuddy — vincent_koc · 2026-08-28
- Midjourney prompting tip: using 'imperfection' to narrate the aftermath of luxury — tisch_eins · 2026-08-28