OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
Zhongyu Yang, Jiale Tao, Ruitao Chen, Zuhao Yang, Yingfang Yuan, Xueliang Zhao, Auden, Kai Wang, Shuai Shao, Biao Wang, Steve Yves, Qinglin Lu
NeurIPS 2026
cs.CV
2026-10-09
On 786 videos, atomic-unit scoring splits caption pairs tied near 50/50 into 74.2 versus 25.8. Cross-shot identity peaks at 37.81%.
Audio-visual caption scores still come from one of two units. A judge model reads the whole caption and returns one number. The number covers the clip, and it hides a swapped person, a slipped timestamp, or a sound pinned to the wrong shot. Probe sets ask a few cloze or QA questions. Those questions can land on a local mistake, but they sample the video thinly, and two judges often flip the same probe. LongVALE and TimeChat parse free text into an event list before scoring, so a bad parse becomes a bad score.
The unit and the scorer are coupled. When the unit is a paragraph, one model call has to extract, align, and compare, and wording gets mixed with facts. OmniCapBench, from Tencent Hunyuan on the NeurIPS 2026 evaluations and datasets track, changes the target. The model emits atomic units, one entity, shot, or sound at a time. Rules check structure. A language model only compares short fields that already line up.
The schema extends Hunyuan's Multi-Stream Scene Script: separate streams for entities, audio, and picture. Shots split further into subshots, and timestamps become continuous ranges. References give stable IDs to people, objects, and scenes. Events hold sounds and dialogue with a start and an end. Shots cut the visual timeline and attach those people and sounds to a shot, which is the clock the other tracks use.
Clips must contain audio, stay within five minutes, and clear a per-minute density bar for shots and events. Annotation order is fixed: references, then events, then shots. A loop generates, validates, rewrites wording, and selects. IDs, time bounds, and cross-links stay frozen during the rewrite, so polishing cannot invent a new link. A red timestamp burned into the corner is an annotation aid. Models under test see the original frames.
The model returns the three tracks as JSON, at temperature 0, with a 4096-token cap. Invalid JSON fails the schema gate, SGC, and that video scores 0 on every later metric. Rules then drop undefined IDs, reversed times, and event-shot links that do not overlap in time. Match bars are fixed: shot temporal IoU at least 0.30, non-dialogue events at least 0.20, dialogue at least 0.50 on one minus word error rate. Identity and subshot matching compare only local appearance text. The LLM then judges semantic equivalence on fields that already matched.
One action can be a single subshot in the reference and three in the prediction, with both factually fine. References and subshots are scored in both directions. Recall penalizes omissions, precision penalizes inventions, and F1 is the harmonic mean.
The set has 786 videos and 12.8 hours, every clip with audio, test only. Median length is 28.32 seconds and the mean is 58.82. 482 clips, 61.32%, are under a minute; 228 run from 1 to 3 minutes; 76 run from 3 to 5. A video averages 7.40 entities, 14.53 shots, 49.82 subshots, and 8.32 audio events, 6.83 of them dialogue. There are 20 top-level categories and 125 second-level ones.
Rule scores are macro-averaged over videos. Figures below are percentages.
| Model | SGC | CCC | Ref Subject F1 | Event F1 | EVSA | Speaker F1 |
| Gemini 3.1-Pro | 97.36 | 34.43 | 78.60 | 63.32 | 51.46 | 91.18 |
| Gemini 2.5-Pro | 96.73 | 37.81 | 80.70 | 58.52 | 43.74 | 89.72 |
| Qwen3.5-Omni-Plus | 96.09 | 29.55 | 73.57 | 53.92 | 40.20 | 91.20 |
| Qwen3.5-Omni-Flash | 95.58 | 24.31 | 72.27 | 50.68 | 35.58 | 89.91 |
| Seed2.0 | 87.68 | 31.88 | 65.29 | 4.54 | 3.52 | 91.80 |
| MiMo-2.5 | 96.58 | 23.91 | 61.19 | 43.82 | 31.69 | 86.62 |
Schema compliance is mostly solved for the proprietary APIs. Format errors are 1.3% of Gemini 3.1-Pro's failure mass. Identity is not. Gemini 2.5-Pro reaches 80.70 Ref Subject F1 overall and 84.80 on clips under a minute, then falls to 37.81 CCC. Gemini 3.1-Pro is lower on both, at 78.60 and 34.43. Qwen3-Omni-Instruct, Qwen3-Omni-Captioner, and MiniCPM-o-2.6 land at 9.85, 11.51, and 4.29 CCC, while their SGC scores are still 89.02, 90.77, and 72.13. The person is visible in one shot. After the cut, the ID breaks.
Attaching a sound to the shot on screen is harder than detecting the sound. Gemini 3.1-Pro drops from 63.32 Event F1 to 51.46 EVSA. Open-source EVSA falls to 13.33, 17.42, and 2.46, all under 20. Dialogue semantics sit at or above 80%. Seed2.0 scores 91.80 Speaker F1 with an Event F1 of 4.54. The speaker number covers the dialogue lines it actually wrote.
Fine boundaries do not separate the proprietary models. Subshot tIoU sits between 46.51 and 49.18 for all six. Shot-level recovery still does: Gemini 3.1-Pro Shot F1 is 76.24, Gemini 2.5-Pro 70.93.
Semantic scores use only units that already matched. Gemini 3.1-Pro is conservative on people, precision 67.07 and recall 43.78. MiniCPM-o-2.6 posts 76.32 scene precision and 20.39 scene recall. Of Gemini 3.1-Pro's failures, action hallucination is 28.5%, audio-visual mis-attachment 23.7%, and audio hallucination 17.9%.
Judge choice moves holistic scores and barely moves this benchmark. On the same outputs, Qwen3.6-27B, GPT-4o, and Gemini-2.5-Pro score a whole-caption protocol at 71.24, 83.08, and 93.92. QA-style Omni-Cloze scores 61.17, 53.49, and 56.57. The stronger judge is not the higher one. With the LLM confined to local semantics, OmniCapBench scores 56.82, 58.24, and 57.60.
Whole-caption metrics also flatten real gaps. On 200 videos from UGC-VideoCap and video-SALMONN 2, the pairs Qwen3.5-Omni-Flash versus Qwen3-Omni, and Plus versus Flash, were kept only where text metrics sat near 50 versus 50. Under structural checks the stronger side reaches 74.2, in line with a human audit, and the weaker side falls to 25.8. In one traced case, Qwen3.5-Omni-Plus invents protective eyewear on a person using a leaf blower, and the paragraph still reads cleanly. ASID-Caption and AVoCaDO stay out of the main table. Supervised fine-tuning on free-form captions left them unable to follow the JSON instruction.
For proprietary models, schema compliance is already high. The spread sits in cross-shot identity and in whether a sound is attached to the picture that is playing. A single caption score keeps ranking fluent paragraphs above models that hold an ID. The 74.2 to 25.8 split is that gap, measured where text scores claimed a tie.
Using the benchmark requires the three-track JSON. A zero from a failed SGC says nothing about what the model saw. Inside the main table, format failures are a small slice of the error, so the remaining spread is mostly perceptual. Nothing here shows that training on this schema lifts CCC.
The abstract calls the weakness long-horizon. The median clip is 28.32 seconds, 61% are under a minute, and nothing runs past five minutes. Identity already breaks at that length. That is a real defect. It is not a measurement of video past ten minutes.
The paper does not report inter-annotator agreement. References come from a multi-model draft plus human revision. Where a subshot splits, and whether a noise counts as its own event, can follow the generators' habits. Nobody measures whether models from the same family are favored.
Local semantic judgments still use an LLM. Flat totals across three judges mean a stronger judge no longer inflates the score. They do not mean each local call matches a person. The 74.2 that lines up with a human audit comes from those 200 videos preselected for a text-score tie, not from a random draw of the 786.
Statements outside the annotation budget do not directly cut the main precision. True details the annotators skipped go unpunished. Some unsupported extras miss the main score the same way.
A 4096-token cap, on videos that average about 50 subshots, mixes truncation into recall. The paper does not separate 'did not see it' from 'could not fit it'. The match bars have no sensitivity check. Non-dialogue events match at a tIoU of 0.20, so Gemini 3.1-Pro can show Event tIoU 71.92 beside Event F1 63.32. Macro-averaging gives the 482 short clips the most weight. How CCC and EVSA move with duration is not tabled.