NJU Introduces AVE-Compass: A Benchmark for Audio-Video Editing
NJU-LINK · hf · 2026-08-06
A team from Nanjing University introduced AVE-Compass, a benchmark designed to evaluate complex audio-visually coupled editing. While existing benchmarks often evaluate audio and visual modalities in isolation, AVE-Compass features 145 source videos, 196 coupled editing instructions, and 2,688 fine-grained checklist items. It assesses models across instruction following, fidelity preserving, realism, and editing intent.
Evaluations reveal that state-of-the-art models still struggle with cross-modal instructions while preserving non-target content. To address this, the researchers proposed AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively refines results via self-reflection and evaluator feedback, effectively improving instruction execution and audio-visual alignment.
More from Multimodal
- Seedance 2.5 in Action: Generating Emotional Video Memories with AI — eyishazyer · 2026-08-06
- Hands-on: MiniMax H3 Lags Behind Veo 3 and Seedance in Video Generation — yolaoheinz · 2026-08-06
- Midjourney Guide: Creating Retro Dotted Painting Style with Dual Moodboards — _AustinCalvert_ · 2026-08-06
- Trying to Make a 4-Minute Video with Local AI Setup — Then-Comfortable8258 · 2026-08-06
- Generating Found-Footage Fantasy Short Films with H3: Complex Prompt Test — Disastrous-Agency675 · 2026-08-06
- WISRD Benchmark: Evaluating AI Problem-Solving in Pure Image Space — HirokatuKataoka · 2026-08-06