Video-DeepResearch Outperforms GPT-5 and Claude in Multimodal Agents
Zhen Fang · hf · 2026-08-05
This paper introduces Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, which demands dense spatiotemporal grounding coupled with open-web exploration.
- Bottlenecks Addressed: Overcomes modality bias (agents bypassing visual tools for text) and parametric knowledge leakage (relying on memory instead of tools).
- Methodology: Features a decoupled perception-exploration pipeline with stage-wise tool unlocking to force cross-frame visual grounding. Trained via SFT followed by GRPO.
- New Benchmark: Curates Video-DR-Bench, comprising 200 complex, multi-hop VQA instances.
- Results: Video-DeepResearch-35B-A3B achieves 64.0% average accuracy (SOTA), surpassing Claude-4.5-Sonnet (59.0%), GPT-5 (52.5%), and Gemini 2.5 Pro (57.5%).
More from coding & agent
- Developer Waited 6 Months for Claude, Built Entire App in One Day — airkatakana · 2026-08-05
- Proposed cmux Multi-Column Sidebar: Machines, Workspaces, and Agents — philipvollet · 2026-08-05
- Chinese Models Dominate Forecasting Leaderboard with Advanced AI Agents — teortaxesTex · 2026-08-05
- eidoverse-worlds: Open-Source Persistent 3D World Where Humans and AIs Co-Exist — repligate · 2026-08-05
- AARM Spec for AI Agent Audit Trails Released: 9 Properties to Fight Memory Poisoning — Funky_Chicken_22 · 2026-08-05
- What is a 'Software Factory'? AI Agents Reshape the SDLC — Pavan_Belagatti · 2026-08-05