VideoCoCo: Enhancing Video Physical Consistency via Blender Code-as-CoT

Haodong Li · hf · 2026-07-31

While text-to-video models achieve stunning visual quality, they often struggle with physically consistent dynamics. This paper introduces VideoCoCo, an agentic dual-engine framework where executable Blender code serves as a process-level chain of thought.

The workflow includes:

This approach separates process-level reasoning from visual realization, significantly improving baseline scores on PhyGenBench and VBench-2.0.

Original post →

More from Multimodal

Multimodal channel →