TRACE benchmark shows near-identical QA scores mask big gaps in streaming video models

omlab · hf · 2026-09-28

omlab released TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark for streaming video understanding that makes evidence timing, history maintenance, and response triggering explicit.

Why: current evaluations report task scores without specifying when evidence becomes valid or how responses are triggered, so similar scores can hide very different workloads and failure modes.

How: TRACE combines temporally audited tasks with evidence-timing and trigger annotations, a causal Core–Adapter protocol that controls information availability while logging actual history processing and response events, and multidimensional reporting (answer quality, timeliness, response-selection behavior, workload, completion, reliability).

Findings: across 1,240 records from 517 videos and 8 models in 8 configurations, nearly identical QA accuracy masked large differences in completion, answer validity, and generation workload. Proactive performance further splits into response quality, delay, false alarms, and missed target windows. Code and benchmark are open-sourced on GitHub.

Original post →

More from Multimodal

Multimodal channel →