New Benchmark for Multi-Reference Audio-Visual Generation
KlingTeam · hf · 2026-07-17
This work proposes MultiRef-Compass to evaluate Multi-Reference Audio-Visual Generation (MR2AV).
Motivation
Existing benchmarks mostly focus on:
- Text-driven generation
- Single-reference subject preservation
- Consistency in only a part of image or audio/video
However, MR2AV requires models to simultaneously handle multiple references and generate synchronized audio-visual content, making it much harder.
Benchmark Design
The authors built a unified benchmark of 350 samples, covering:
- Multi-view subject preservation
- Multi-entity binding
- Person-object-scene composition
Evaluation Protocol
Defined 4 evaluation dimensions with 14 sub-metrics:
- Basic quality
- Reference consistency
- Audio-visual consistency
- Instruction following
Adopts an "automatic metrics + rejudging enhanced MLLM-as-a-Judge" framework, making the evaluation more scalable and auditable.
Experimental Conclusions
Across 8 representative MR2AV systems, the authors found current methods still have significant room for improvement across multiple dimensions, indicating this direction is far from solved.
More from Multimodal
- Seedance 2.0 demo turns ketchup on spaghetti in Rome into an AI reaction meme — azed_ai · 2026-07-21
- A reusable “Lunar Eclipse Dreamscape” prompt comes with multiple example renders — LudovicCreator · 2026-07-21
- Midjourney 8.2 preview shows a double-exposure prompt with strong style control — michaelrabone · 2026-07-21
- Travel MCP Server adds flight, hotel, weather and budget tools for agents — modelcontextprotocol · 2026-07-21
- Douyin Video Analysis MCP turns share links into structured video summaries — modelcontextprotocol · 2026-07-21
- Synthesia launches Dubbing 2.0 with 130+ languages and lip-sync video translation — synthesiaIO · 2026-07-21