OmniCapBench: 786-video benchmark exposes weak long-horizon audio-visual reasoning
Tencent-Hunyuan · hf · 2026-10-09
Tencent Hunyuan released OmniCapBench, a deep-structured evaluation framework for fine-grained audio-visual captioning by MLLMs.
- Current benchmarks trade off coverage vs localization, and unconstrained LLM judges are unstable; OmniCapBench instead targets atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, scored via deterministic constraint checks plus localized LLM semantic comparison.
- The benchmark contains 786 densely annotated videos and separates perception errors like temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions.
- Frontier MLLMs show strong local perception but weak long-horizon audio-visual reasoning, especially identity drift and cross-modal misalignment, offering a fine-grained roadmap for omnimodal development.
More from Multimodal
- Scanography-style portraits on Midjourney v8.2, full prompt included — tisch_eins · 2026-10-09
- Both influencers are AI: Higgsfield Katana makes synthetic creators trivial — CurieuxExplorer · 2026-10-09
- Claude Code Drives ComfyUI with Krea2 and LTX 2.5 in Hilarious AI Video Test — TheHollywoodGeek · 2026-10-09
- 20 GitHub repos turn Claude into a motion design studio — Roger_M_Taylor · 2026-10-09
- Tencent Hunyuan's training-free MC-Sparse attention speeds DiT denoising up to 2.3x — Tencent-Hunyuan · 2026-10-09
- Setting camera paths freely inside generated scenes is a surprisingly cool experience — XRarchitect · 2026-10-09