Salesforce Introduces Traceable Evidence Video QA
Salesforce · hf · 2026-07-15
Salesforce proposed Evidence-Backed Video Question Answering (E-VQA), requiring video models to output verifiable visual evidence alongside their answers: temporal segments and densely tracked object segmentation masklets.
They also released ST-Evidence, the first human-verified benchmark covering both discriminative and generative tasks with pixel-level grounding. Experiments reveal a clear disconnect between the QA accuracy of current Video LLMs and their actual visual perception abilities—a gap that cannot be bridged simply by scaling up model size.
To address this, the team built a scalable automated data generation pipeline to create ST-Evidence-Instruct, a 160,000-instance dataset bridging high-level reasoning and fine-grained grounding. Fine-tuning grounded Video LLMs with this data yielded significant improvements. For instance, a 7B model outperformed the similarly sized UniPixel baseline, boosting t-mean by 27.2 and J&F by 13.8. Code and data have been open-sourced.
More from Research
- SUFLECA shows NOC-based correspondence can improve CAD-to-image alignment — ducha_aiki · 2026-07-21
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21