CoEvoWhen co-evolves policies and tools for ultra-long video temporal grounding

zju · hf · 2026-10-01

Zhejiang University's CoEvoWhen jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories, forming a reusable skill without updating model parameters.

An external skill updater distills transferable long-video task experience, refining orchestration of long-range image-based and fine-grained video-based observations, while using coding capabilities to upgrade or create tools. The evolved VLM autonomously orchestrates tools without a stronger separate planner. Across five benchmarks and three VLMs, co-evolution consistently improves ultra-long video grounding accuracy while reducing visual token cost, and the evolved skill transfers to general long-video QA with no task-specific evolution.

Original post →

More from Research

Research channel →