Salesforce Introduces Traceable Evidence Video QA

Salesforce · hf · 2026-07-15

Salesforce proposed Evidence-Backed Video Question Answering (E-VQA), requiring video models to output verifiable visual evidence alongside their answers: temporal segments and densely tracked object segmentation masklets.

They also released ST-Evidence, the first human-verified benchmark covering both discriminative and generative tasks with pixel-level grounding. Experiments reveal a clear disconnect between the QA accuracy of current Video LLMs and their actual visual perception abilities—a gap that cannot be bridged simply by scaling up model size.

To address this, the team built a scalable automated data generation pipeline to create ST-Evidence-Instruct, a 160,000-instance dataset bridging high-level reasoning and fine-grained grounding. Fine-tuning grounded Video LLMs with this data yielded significant improvements. For instance, a 7B model outperformed the similarly sized UniPixel baseline, boosting t-mean by 27.2 and J&F by 13.8. Code and data have been open-sourced.

Original post →

More from Research

Research channel →