Technion's MIST stress test finds irrelevant images shift 20% of VLM judge labels regardless of content
Technion · hf · 2026-10-01
Technion researchers introduce MIST (Misleading-Image Stress Test): 200 English sentences readable figuratively or literally, paired with an aligned image, a misleading one, or none, with guidelines requiring the label to be decided from the sentence alone.
Key findings:
- Across 13 VLM judges, an aligned image flipped 20.5% of labels and a misleading one 19.4% — nearly identical, and both above the 11.6% caused by simply deleting the ignore-the-image instruction.
- Only 37% of flipped labels moved toward the sense the image actually depicted; agreement with human annotators was unchanged whether images were absent, aligned, or misleading.
- What moves a judge is that an image is present, not which one it is — so substitutability verdicts for VLM-as-annotator describe the evaluation configuration as much as the model itself.
More from Models
- GPT-6 Pro weekly message cap reportedly cut from 200 to 100 — koltregaskes · 2026-10-01
- Grok Bot Gains 'Primary Bot' That Manages Other Bots Proactively in Latest iOS App — testingcatalog · 2026-10-01
- Pretraining 800 LMs shows AI-generated web text can actively hurt scaling — iScienceLuvr · 2026-10-01
- Endless Exam benchmark tests 9 models on 14 families of mathematical constructions beyond human frontiers — Muhan Zhang · 2026-10-01
- Gemini 4 Argon jumps to 57.6% on Terminal-Bench-Science but underperforms on TB 4.0 — JJitsev · 2026-10-01
- Hume AI CEO and CPO on why voice models still can't truly grasp tone and emotion — kimmonismus · 2026-10-01