CoEvoWhen co-evolves policies and tools for ultra-long video temporal grounding
zju · hf · 2026-10-01
Zhejiang University's CoEvoWhen jointly evolves high-level policies and executable media tools from a VLM's agentic reasoning trajectories, forming a reusable skill without updating model parameters.
An external skill updater distills transferable long-video task experience, refining orchestration of long-range image-based and fine-grained video-based observations, while using coding capabilities to upgrade or create tools. The evolved VLM autonomously orchestrates tools without a stronger separate planner. Across five benchmarks and three VLMs, co-evolution consistently improves ultra-long video grounding accuracy while reducing visual token cost, and the evolved skill transfers to general long-video QA with no task-specific evolution.
More from Research
- ArchMap lands in Nature Genetics: code-free single-cell mapping onto reference atlases — burny_tech · 2026-10-01
- New Paper Argues Literary Tools Are Essential for Building Culturally Literate AI — begusgasper · 2026-10-01
- NVIDIA's Instant NuRec reconstructs a drivable 3DGS world from driving logs in ~1.5 seconds — rsasaki0109 · 2026-10-01
- KV-streams trains SWE agents 2x faster by preserving KV cache across compaction — burny_tech · 2026-10-01
- Researcher argues "ego" beats "persona" for describing LLM identity — repligate · 2026-10-01
- Silicon microring modulators push past 200Gb/s per lane to cut AI optical I/O power — jwt0625 · 2026-10-01