CLIP-CC-Bench: A New Benchmark for Evaluating Paragraph-Level Video Descriptions
MINT-SDSU · hf · 2026-08-11
Current evaluations of video-language models mostly focus on short clips and single-sentence metrics. To address the gap in assessing long-form, paragraph-level descriptions, researchers introduced CLIP-CC-Bench.
Key features and methodology include:
- Data Construction: Built on 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference.
- Evaluation Methodology: Employs an ensemble of five state-of-the-art LLM-based embedding models to mitigate single-model bias. It compares model-generated descriptions against references using coarse-grained and fine-grained semantic matching.
- Results: Evaluated 17 state-of-the-art video-language models, reporting Borda-aggregated rankings and average scores. The protocol's internal reliability was verified through inter-judge agreement and bootstrap ranking stability.
Standardized evaluation scripts, model outputs, and aggregation tools have been open-sourced to support reproducibility.
More from Multimodal
- Minimax H3 Model Impresses in Tests: Breaks 15-Second Limit with High Quality and Coherence — tsi_org · 2026-08-11
- Testing MiniMax Music Model: Mimicking Artists via Text Prompts — elfpresetsv3 · 2026-08-11
- Suno's watermarking criticized: may drive traditional musicians away — koltregaskes · 2026-08-11
- RTX 5090 vs. Dual 48GB GPUs: A Hardware Upgrade Guide for Local AI Video Generation — Ammoryyy · 2026-08-11
- Crafting psychological horror with Dreamina: less sound is more — Eric520CC · 2026-08-11
- Troubleshooting ComfyUI: Why T2V Generation is Slower Than Ref2V — 5tephaniehemming · 2026-08-11