MMOOC Benchmark: Multimodal LLMs Struggle to Balance In-Context and Out-of-Context Questions
VCLab-HKPU · hf · 2026-08-11
MMOOC is a large-scale benchmark designed for out-of-context evaluation in Multimodal Large Language Models (MLLMs).
It assesses whether models can correctly refuse out-of-context questions while answering shifted in-context questions. The research reveals that current models struggle significantly to balance these two abilities.
More from Models
- Muse Glimmer 30B Coding Test: Zero Invented Defects in Working Code — pbaylies · 2026-08-11
- NVIDIA Nemotron 3.5 Lightning Hits DeepInfra with 1M Token Context — gharik · 2026-08-11
- Upstage Launches Solar Pro 4 for Production Agents, Free on Nous Portal — NousResearch · 2026-08-11
- Nemotron 3.5 Lightning Tested: 670 tokens/s Speedster for Agents — ArtificialAnlys · 2026-08-11
- OpenAI Pauses High-Risk Astra but Ships GPT-5.6-Cyber — eyishazyer · 2026-08-11
- ExtractBench: Commercial VLM Recall Drops Below 35% on 50+ Page Enterprise Docs — llama_index · 2026-08-11