C4 Benchmark: MLLMs Struggle with Cross-Concept Understanding and Creative Decoding

kinamind · hf · 2026-08-10

Evaluating the creative capabilities of Multimodal LLMs (MLLMs) is challenging due to the lack of explicit targets and reward signals. Researchers introduce C4, a cognition-inspired evaluation framework that tests cross-concept creativity by measuring the ability to understand implicit conceptual relations in Chinese idioms (Chengyu).

The study features the C4-Eval set with 184 synthetic items and 37 human-created figures. Across ten evaluated MLLMs, the strongest closed models achieved only 50.7% accuracy, with open-source models performing substantially lower. While candidate constraints improved accuracy sharply, bridge hints and explanation requests provided modest gains, exposing a significant gap in how current MLLMs decode creatively encoded meaning.

Original post →

More from Research

Research channel →