VideoCoCo: Using Executable Code to Fix Video Physics

jiqizhixin · x · 2026-08-23

CUHK-Shenzhen and USTC introduce VideoCoCo to fix physics issues in text-to-video models. Instead of guessing physics in pixel space, it generates executable code describing the physical process step-by-step as a 'code-as-chain-of-thought'. A dual-engine agent then uses this code to generate the video. On PhyGenBench, it boosts the average score from 0.475 to 0.558, beating the Wan2.2-TI2V-5B baseline and ranking first in mechanics, optics, and materials. On VBench-2.0, the average score jumps from 52.18% to over 77%.

Original post →

More from Multimodal

Multimodal channel →