VideoCoCo: Enhancing Video Physical Consistency via Blender Code-as-CoT
Haodong Li · hf · 2026-07-31
While text-to-video models achieve stunning visual quality, they often struggle with physically consistent dynamics. This paper introduces VideoCoCo, an agentic dual-engine framework where executable Blender code serves as a process-level chain of thought.
The workflow includes:
- Code Generation: A coding agent synthesizes a Blender program specifying the scene and temporal evolution.
- Simulation: The executable engine runs the program to produce a deterministic spatiotemporal draft.
- Visual Rendering: A generative video engine transforms the draft into a photorealistic video.
This approach separates process-level reasoning from visual realization, significantly improving baseline scores on PhyGenBench and VBench-2.0.
More from Multimodal
- H3-Omni Architecture: Native 2K Resolution and In-Context Regeneration — bdsqlsz · 2026-07-31
- Flux 3 and Minimax H3 Announce Open Weights for Video Generation — cocktailpeanut · 2026-07-31
- Runway Video Test: Generating a 15-Second Mumbai Monsoon Timelapse — CurieuxExplorer · 2026-07-31
- Testing Reference Image Generation with Hailuo-03 — chrisfirst · 2026-07-31
- MiniMax-H3 Tops Video Editing Leaderboard, Ties for 2nd in Text-to-Video — bdsqlsz · 2026-07-31
- Generating GPU ASMR: Flux 3 Opens Early Access on Hermes Agent — venturetwins · 2026-07-31