NJU Introduces AVE-Compass: A Benchmark for Audio-Video Editing

NJU-LINK · hf · 2026-08-06

A team from Nanjing University introduced AVE-Compass, a benchmark designed to evaluate complex audio-visually coupled editing. While existing benchmarks often evaluate audio and visual modalities in isolation, AVE-Compass features 145 source videos, 196 coupled editing instructions, and 2,688 fine-grained checklist items. It assesses models across instruction following, fidelity preserving, realism, and editing intent.

Evaluations reveal that state-of-the-art models still struggle with cross-modal instructions while preserving non-target content. To address this, the researchers proposed AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively refines results via self-reflection and evaluator feedback, effectively improving instruction execution and audio-visual alignment.

Original post →

More from Multimodal

Multimodal channel →