CLBench-V: New multimodal context learning benchmark shows top score below 0.3

Lai Wei · hf · 2026-07-30

While existing evaluations focus mainly on text, real-world tasks often require models to learn from multimodal contexts like figures and web pages. To address this, researchers introduced CLBench-V, a benchmark organizing tasks across three dimensions: context grounding, new information application, and new knowledge learning.

Spanning domains like science, finance, and spatial reasoning with 3,443 instances, the evaluation of six recent multimodal models reveals that the best overall score is only 0.2847, indicating the capability is far from saturated. InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus leads in new information application.

Original post →

More from Models

Models channel →