MuseBench: Evaluating MLLMs' Intent Understanding in Audiovisual Art

nanyang-technological-university-singapore · hf · 2026-07-08

Nanyang Technological University released MuseBench, a comprehensive benchmark designed to evaluate multimodal large language models' ability to understand audiovisual artworks at the intent level. Evaluations reveal a significant gap between current mainstream models and human experts in professional creative understanding, with models performing particularly poorly in scenarios requiring deep cultural context and intent interpretation.

Related event: NTU Introduces MuseBench for Multimodal Art Understanding(2 posts)→

Original post →

More from Multimodal

Multimodal channel →