VideoCoCo: Using Executable Code to Fix Video Physics
jiqizhixin · x · 2026-08-23
CUHK-Shenzhen and USTC introduce VideoCoCo to fix physics issues in text-to-video models. Instead of guessing physics in pixel space, it generates executable code describing the physical process step-by-step as a 'code-as-chain-of-thought'. A dual-engine agent then uses this code to generate the video. On PhyGenBench, it boosts the average score from 0.475 to 0.558, beating the Wan2.2-TI2V-5B baseline and ranking first in mechanics, optics, and materials. On VBench-2.0, the average score jumps from 52.18% to over 77%.
More from Multimodal
- Top YouTube creators face backlash over paid Higgsfield/Seedance demos built on real people's likenesses — emmanuelvivier · 2026-08-23
- Minimax H3 pruned 20B at 832x480, upscaled 2x to 1664x960 via LTX 2.3 — call-lee-free · 2026-08-23
- Minimax H3 Tuning: Sampler and Scheduler Selection Guide — LFAdvice7984 · 2026-08-23
- DeepSeek-V4-Flash-Vision-EXP Generation Demo Shared — op7418 · 2026-08-23
- Minimax H3 Model Generates Consistent Character Sheets and Turnarounds — PoopMan333 · 2026-08-23
- Minimax H3 Image-to-Video Demo: Lisbon Finally Snaps — MikePounce · 2026-08-23