TRACE benchmark shows near-identical QA scores mask big gaps in streaming video models
omlab · hf · 2026-09-28
omlab released TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark for streaming video understanding that makes evidence timing, history maintenance, and response triggering explicit.
Why: current evaluations report task scores without specifying when evidence becomes valid or how responses are triggered, so similar scores can hide very different workloads and failure modes.
How: TRACE combines temporally audited tasks with evidence-timing and trigger annotations, a causal Core–Adapter protocol that controls information availability while logging actual history processing and response events, and multidimensional reporting (answer quality, timeliness, response-selection behavior, workload, completion, reliability).
Findings: across 1,240 records from 517 videos and 8 models in 8 configurations, nearly identical QA accuracy masked large differences in completion, answer validity, and generation workload. Proactive performance further splits into response quality, delay, false alarms, and missed target windows. Code and benchmark are open-sourced on GitHub.
More from Multimodal
- GPT-6 Astra demo turns real-room video into interactive 3D worlds for robot training — 141_1337 · 2026-09-28
- Quantum Neon Origami: a glowing geometric-fold prompt template for image generation — LudovicCreator · 2026-09-28
- Midjourney --raw parameter: turning off the auto-pilot for photorealistic images — michaelrabone · 2026-09-28
- SparkDiffusion from PKU/Tsinghua/Alibaba hits 265x video generation speedup on a single RTX 5090 — 机器之心 · 2026-09-28
- Higgsfield animation team shares full AI pipeline: 18 shots of a car chase in 3 days — xiaohu · 2026-09-28
- First open-weight adult video generation model lands on Hugging Face — jedisct1 · 2026-09-28