New Benchmark for Multi-Reference Audio-Visual Generation

KlingTeam · hf · 2026-07-17

This work proposes MultiRef-Compass to evaluate Multi-Reference Audio-Visual Generation (MR2AV).

Motivation

Existing benchmarks mostly focus on:

However, MR2AV requires models to simultaneously handle multiple references and generate synchronized audio-visual content, making it much harder.

Benchmark Design

The authors built a unified benchmark of 350 samples, covering:

Evaluation Protocol

Defined 4 evaluation dimensions with 14 sub-metrics:

Adopts an "automatic metrics + rejudging enhanced MLLM-as-a-Judge" framework, making the evaluation more scalable and auditable.

Experimental Conclusions

Across 8 representative MR2AV systems, the authors found current methods still have significant room for improvement across multiple dimensions, indicating this direction is far from solved.

Original post →

More from Multimodal

Multimodal channel →