VIABench Evaluates Visual Assistance Capabilities

NJU · hf · 2026-07-17

The research team introduces VIABench, a comprehensive video benchmark designed for visual assistance scenarios, featuring first-person perspective videos recorded or shared by visually impaired individuals. It specifically evaluates the capabilities of multimodal large models in real-world assistive tasks, covering three core categories:

The authors also designed a rigorous evaluation pipeline supporting both online and offline modes. Experimental results show that current multimodal large models still exhibit significant shortcomings in real-world visual assistance, particularly in proactive alerting tasks requiring anticipation and real-time responses. Code and data will be open-sourced.

Original post →

More from Multimodal

Multimodal channel →