FATE Model: Achieving Dual Frame-Level Semantic and Temporal Alignment for Audio-Visual

RUC · hf · 2026-08-10

Current audio-visual models often struggle to balance semantic matching with temporal alignment. To bridge this gap, researchers introduced FATE (Frame-level Audio-visual Temporal Embedding).

The source code for the project has been open-sourced.

Original post →

More from Multimodal

Multimodal channel →