OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
Hongyu Li, Manyuan Zhang, Kaituo Feng, Shu Chen, Dian Zheng, Hao Li, Hao Yu, Zhangquan Chen, Zoey Guo, Ray Zhang, Shaofei Huang, Tianrui Hui, Linjiang Huang, Si Liu
cs.CV
2026-10-09
OneSearch-VL trains an 8B agent with a visually grounded evidence graph, beating tool-using Qwen3-VL-8B by 20.2 points on multi-image research and 17.6 on video.
Cropping a region, lining up an entity across several images, and finding the right moment in a video are different visual skills. Deep research here means a multi-turn tool loop: ground a clue in the picture, retrieve a webpage, and compose source-backed facts into an answer.
The workflow is shared. The records usually are not. Perception-to-knowledge chains, factual rubrics, and generic evidence graphs do not keep the full link from a box or frame, to an entity, to a sourced fact, to the operation that builds the answer. Video pipelines often use an entity or a keyframe as a search seed and drop the provenance once the question exists. Failures in localization, retrieval, and composition then cannot be separated.
OneSearch-VL trains one policy for all three input types. The lead group is at BUAA, with coauthors at CUHK, NTU, THU, and HFUT. Question writing, trajectory filtering, process rewards, and operation-level eval all hang off one task-level reference.
That reference is the Visually Grounded Evidence Graph (VGEG). A video is first an input-level graph of events, keyframes, boxes, and web entities. Each question is a projection: the anchors it needs, the source-backed facts, and the operation that turns those facts into an answer. Single-image data are not recollected. Wikipedia entity paths from OpenSearch-VL are converted into the same graph.
The video pool starts at 2.5 million YouTube records and ends at 70,781. Filters cover upload date (2023 onward), metadata, category quotas, duration balance, and a GPT-4.1 pass on titles and descriptions. Music and education are removed because searchable entities are scarce. Frames are sampled at 2 FPS in clips of ten. Seed 2.0 Pro writes descriptions and proposes entities. Qwen3-VL-Embedding drops near-duplicate frames when image similarity is at least 0.9 or text similarity is at least 0.8. Clips are grouped into events, Qwen3-VL-30B boxes objects, and Qwen3-32B checks webpages. A box goes to reverse image search, or to OCR followed by text search. Each fact stores a head, a relation, a tail, a URL, and a quotation, then expands through successors scored at least 0.5.
A question is kept only if the selected evidence can regenerate the answer. Explicit names are rewritten into descriptions that only the picture can resolve, and a text-only probe deletes items solvable without vision. Multi-image items are shuffled, deduplicated keyframes, with frame ids remapped to image indices.
Seed 2.0 Pro also rolls out the expert trajectories in the real tool environment, without the reference answer. GPT-4o judges the final answer. A process check requires at least one effective tool call. Kept trajectories average 10.1 tool turns. OneSearch-VL-SFT-110K has about 110k trajectories: 36k single-image, 37k multi-image, and 35k video. OneSearch-VL-RL-10K has about 10k tasks: 3.7k, 2.8k, and 3.6k. The base model samples eight times per task, and only tasks answered correctly on 1 to 7 of those runs are kept.
One call protocol covers the tools. Vision tools are crop, OCR, perspective correction, super-resolution, and sharpening. Retrieval tools are image search and text search. Video adds timespan selection and frame extraction. SFT starts from Qwen3-VL-8B, fully updates the language model, and freezes the vision encoder and projector. The loss covers reasoning, tool commands, and the final answer, not tool observations. RL uses GRPO, which updates from the relative rewards inside a group of trajectories for the same question.
The reward is a format-gated sum: answer correctness 0.6, query quality 0.2, and the Evidence-aware Visual-Grounded Rubric reward (EVGR) 0.2. EVGR reads a rubric built from that item's VGEG. One judge asks whether observations support the required facts. The other asks whether the right object, region, or frame was used. Single-image and multi-image inputs are already visible, so an extra visual call is not required. Video must hit the relevant moment and feed that frame's identity into later search. If a tool fails, only the valid prefix is trained.
OneSearch-MI-Bench has 301 questions and OneSearch-Video-Bench has 307, 608 in total. Items are split by the final answer-producing operation, not by topic. Single-anchor lookup is 8.7%. The five compositional operations are 91.3%. Multi-image questions use 2 to 8 images, 3.05 on average. Humans check anchors, sources, and operation labels after the automatic filters.
GPT-4o marks each final answer correct or not, under the same tool setup used in training.
On seven single-image benchmarks, OneSearch-VL-8B averages 58.3. Tool-using Qwen3-VL-8B averages 42.0, 16.3 points lower. OpenSearch-VL-8B averages 56.6, 1.7 lower, and trails on every one of the seven. The paper highlights VDR +2.8 (23.6 vs 20.8), InfoSeek +2.1 (72.3 vs 70.2), and MMSearch +2.0 (66.5 vs 64.5). SenseNova-MARS-8B still leads MMSearch, 67.4 to 66.5. VDR and BrowseComp-VL stay hard at 23.6 and 39.0.
On the new multi-image bench, the new video bench, and VideoDR, the paper reports gains of 20.2, 17.6, and 27.0 points over tool-using Qwen3-VL-8B. The ablation table lists absolute scores of 55.8, 35.5, and 57.0 against 35.5, 17.9, and 30.0. OpenSearch-VL-8B scores 47.0 on VideoDR. Compositional operations move more. On multi-image, knowledge-conditioned counting, multi-anchor arithmetic, and multi-anchor comparison rise by 29.8, 26.3, and 22.5 points. On video, multi-hop retrieval, multi-anchor arithmetic, and multi-anchor join rise by 25.5, 20.0, and 19.2.
| Comparison | Metric | Result |
| vs tool-using Qwen3-VL-8B | 7 image benchmarks, avg | 58.3 vs 42.0 (+16.3) |
| vs OpenSearch-VL-8B | 7 image benchmarks, avg | 58.3 vs 56.6 (+1.7) |
| vs tool-using Qwen3-VL-8B | MI / Video / VideoDR | +20.2 / +17.6 / +27.0 |
| Ablation absolutes | MI / Video / VideoDR | 55.8 / 35.5 / 57.0 |
| Joint SFT, 3 input types | 6-benchmark avg | 55.8; single-type 51.7 to 55.1 |
| Answer + query, then full EVGR | 6-benchmark avg | 57.3 to 61.1 |
The six-benchmark average is SimpleVQA, InfoSeek, FVQA, the two new benches, and VideoDR. It is a different number from the seven-image average. With the same SFT step count, any single input type lifts it from 40.7 into the range 51.7 to 55.1. Video-only is strongest at 55.1, and it raises SimpleVQA to 68.6, above the image-only run at 66.1. The other two single-image scores rise as well. Joint training reaches 55.8, only 0.7 above video-only, and is best in this ablation on SimpleVQA (68.7), multi-image (51.5), and video (31.2).
RL starts at 55.8. Answer-only reward reaches 56.3 and drops multi-image from 51.5 to 50.8. Answer plus query quality reaches 57.3. Traceability alone reaches 59.0. Grounding alone reaches 59.3. Both together reach 61.1, 3.8 above answer-plus-query, best on five benchmarks and tied on the sixth.
On single-image research the step is incremental: +1.7 over OpenSearch-VL, positive on all seven sets, small in size. The training distribution is the part to keep. Multi-image research is its own data, one policy covers all three input types, and video only adds timespan selection and frame extraction. Joint SFT does not give away single-image skill. Video-only training even beats image-only training on SimpleVQA.
Compositional operations are 91.3% of the new benches, and that is where the gains sit. The video bench still ends at 35.5, so finding a moment and then searching is not close to saturated.
Expert trajectories average about 10 tool turns and depend on live image search and web search. SFT used 64 H800s for about four days. RL used 32 H800s for about four days. The vision tower stays frozen. The model learns when to call tools and how to use observations. Without a search API, the 8B weights do not do deep research.
The paper states three limits. TextSearch, ImageSearch, and changing webpages decide whether the evidence is still there, so exact reproduction is not guaranteed. The suggested fixes are retrieval snapshots and a tool-failure test. Automatic VGEG construction and model-based process scores can inject annotation errors and judge bias. Multi-turn tool use is expensive, which calls for adaptive budgets.
Other gaps stay open. Question writing, expert rollouts, and both EVGR judges use Seed 2.0 Pro, while the student is Qwen3-VL-8B. The distillation share is not measured. The two new benches come from the same VGEG pipeline as the training data. Human review checks anchors and sources, yet the distribution still sits closer to the training signal than an external test would. VideoDR, from 30.0 to 57.0, and the seven older image benches are the cleaner external checks. The main table does not include VideoSearcher or Video-DeepResearch. The +27.0 on VideoDR is against tool-using Qwen3-VL-8B and the image agent OpenSearch-VL, which scores 47.0. Joint SFT beats video-only by 0.7, so cross-input complementarity is real and small. The move from 55.8 to 61.1 comes from RL. Difficulty bins are a formula over operation type, hop count, and fact count. The paper says they are not human-calibrated. GPT-4o judges answers. Seed 2.0 Pro judges the process reward. The two judges are never aligned.