ENTRAP-VL: A New Benchmark for Measuring Contextual Entrainment in VLMs
Karan Goyal · hf · 2026-07-24
Researchers introduced ENTRAP-VL, a new benchmark designed to evaluate contextual entrainment in Vision-Language Models (VLMs).
Core Concept: Contextual entrainment is the tendency of a model to let auxiliary input context pull its output, regardless of whether that context is relevant, true, or meaningful. This phenomenon was previously studied only in unimodal language models.
Key Contributions:
- Highlights that in multimodal settings, entrainment becomes a dual phenomenon, drivable independently by textual and visual contexts.
- Introduces a veracity distinction, where context might be generally true in the world but false for the depicted scene.
- Releases a dataset of 1,500 manually curated items across eight categories, organized into textual and visual entrainment streams.
The instrument aims to help the community rigorously investigate contextual biases in VLMs. The dataset will be publicly released.
More from Research
- NeurIPS-era incentives may push AI research toward safer, more repetitive work — serrjoa · 2026-07-24
- Terminal-Bench Team Launches Frontier-Bench for Agent Evaluation — BenBlaiszik · 2026-07-24
- Cisco’s open Antares SLMs tackle vulnerability localization with 350M to 3B parameters — huggingface · 2026-07-24
- BeingBeyond’s Being-M0.7 learns humanoid motion from video and beats prior methods on Unitree G1 — jiqizhixin · 2026-07-24
- Stanford Analyzes 500M Words of Statutes: CA Reporting Requirements Grew 400% — StanfordHAI · 2026-07-24
- NeurIPS Reviews Hit by AI Spam: 60% of Sampled Peer Reviews Suspected ChatGPT-Generated — graph_ · 2026-07-24