New Benchmark for Multi-Reference Audio-Visual Generation
KlingTeam · hf · 2026-07-17
This work proposes MultiRef-Compass to evaluate Multi-Reference Audio-Visual Generation (MR2AV).
Motivation
Existing benchmarks mostly focus on:
- Text-driven generation
- Single-reference subject preservation
- Consistency in only a part of image or audio/video
However, MR2AV requires models to simultaneously handle multiple references and generate synchronized audio-visual content, making it much harder.
Benchmark Design
The authors built a unified benchmark of 350 samples, covering:
- Multi-view subject preservation
- Multi-entity binding
- Person-object-scene composition
Evaluation Protocol
Defined 4 evaluation dimensions with 14 sub-metrics:
- Basic quality
- Reference consistency
- Audio-visual consistency
- Instruction following
Adopts an "automatic metrics + rejudging enhanced MLLM-as-a-Judge" framework, making the evaluation more scalable and auditable.
Experimental Conclusions
Across 8 representative MR2AV systems, the authors found current methods still have significant room for improvement across multiple dimensions, indicating this direction is far from solved.
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11