MMOOC Benchmark: Multimodal LLMs Struggle to Balance In-Context and Out-of-Context Questions

VCLab-HKPU · hf · 2026-08-11

MMOOC is a large-scale benchmark designed for out-of-context evaluation in Multimodal Large Language Models (MLLMs).

It assesses whether models can correctly refuse out-of-context questions while answering shifted in-context questions. The research reveals that current models struggle significantly to balance these two abilities.

Original post →

More from Models

Models channel →