CLBench-V: New multimodal context learning benchmark shows top score below 0.3
Lai Wei · hf · 2026-07-30
While existing evaluations focus mainly on text, real-world tasks often require models to learn from multimodal contexts like figures and web pages. To address this, researchers introduced CLBench-V, a benchmark organizing tasks across three dimensions: context grounding, new information application, and new knowledge learning.
Spanning domains like science, finance, and spatial reasoning with 3,443 instances, the evaluation of six recent multimodal models reveals that the best overall score is only 0.2847, indicating the capability is far from saturated. InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus leads in new information application.
More from Models
- Testing All OpenRouter TTS Models: Kokoro-82M is Best and Cheapest for Long-Form — nathanborror · 2026-07-30
- Qwen3.6 MoE 2-bit Quantized Version Tops Hugging Face Trending — EschaLabs · 2026-07-30
- Leaked Tasks Hint at Anthropic's Strategy: Training Expert Judge Models from Human Traces — burny_tech · 2026-07-30
- Users Praise Grok 4.5 for Blazing Fast Speed and Solid Workhorse Capabilities — XFreeze · 2026-07-30
- Is the 342GB Pruned and Quantized Kimi K3 Usable? — Hannibalj2ca · 2026-07-30
- OpenAI Teases New Codex Update; Community Expects Speed Boosts and Cost Reductions — haider1 · 2026-07-30