Open-Sourcing 14B World Model and 1.3B Version

thetripathi58 · x · 2026-07-10

The post introduces an open-source multimodal world/video generation model, releasing the model weights, inference code, and related agentic-harness code. The main model is 14B, and the paper also introduces a lightweight 1.3B version designed for deployment on a single consumer-grade GPU.

The author also points out a key limitation: the model lacks true long-term memory. When leaving a scene and returning, it will generate a new "plausible-looking" space rather than accurately recalling the previous state.

Related event: Ant-backed Team Open-Sources LingBot-World 2.0 Real-time World Model(26 posts)→

Original post →

More from Multimodal

Multimodal channel →