Study: VLM top-k logits leak far more task-irrelevant image information than expected

A paper by @sineadwilliamso and team, to be presented at COLM 2026, titled "What do your logits know", examines a simple but important question: when you query a vision-language model about an image with yes/no questions, how much task-irrelevant information do its outputs leak? The answer: far more than expected. The finding has direct implications for privacy, fairness, distillation, and interpretability—and should concern both API providers and the research community.

Confirmed

Why It Matters

2026-10-07 ~ 2026-10-07 · 8 related posts

Primary sources