Study: VLM top-k logits leak far more task-irrelevant image information than expected
A paper by @sineadwilliamso and team, to be presented at COLM 2026, titled "What do your logits know", examines a simple but important question: when you query a vision-language model about an image with yes/no questions, how much task-irrelevant information do its outputs leak? The answer: far more than expected. The finding has direct implications for privacy, fairness, distillation, and interpretability—and should concern both API providers and the research community.
Confirmed
- The paper systematically compares information retained at different representation levels of vision-language models: from the information-rich residual stream, through two natural bottlenecks—tuned lens projections and final top-k logits.
- The residual stream encodes nearly everything about a scene, whether or not it's queried; even top-2 logits leak task-irrelevant information—for example, asking "Is there a blue ball?" can expose a ball's material and size that were never mentioned.
- After dimension matching, top-k logits leak roughly as much information as the tuned lens trajectory, while being far more accessible to end users.
- Using only the top-20 logits exposed via the API, about 100 queries suffice to recover attributes that were never asked about.
- The paper concludes that final logits are far from a minimal information bottleneck, with direct implications for privacy, fairness, distillation, and interpretability.
Why It Matters
- Many inference APIs expose top-k logits to end users, meaning ordinary users can extract attributes the model perceives but was never asked about within a small number of queries—a real privacy risk.
- For distillation practitioners, top-k logits are an easier signal source than internal representations; interpretability researchers also need to reassess how much information the final output layer actually contains.
2026-10-07 ~ 2026-10-07 · 8 related posts
Primary sources
- COLM 2026 paper: yes/no queries leak far more via logits than expected — sineadwilliamso ·
- COLM 2026 paper maps how much VLM representations leak across layers — sineadwilliamso ·
- Top-20 logits alone let ~100 queries recover attributes never asked — sineadwilliamso ·
- [source] COLM 2026 paper: yes/no queries leak far more via logits than expected — sineadwilliamso · 2026-10-07
- [source] COLM 2026 paper maps how much VLM representations leak across layers — sineadwilliamso · 2026-10-07
- Ask about a blue sphere and the logits leak its material and size — sineadwilliamso · 2026-10-07
- Target attributes reliably recoverable at k≈10-13, study finds — sineadwilliamso · 2026-10-07
- [source] Top-20 logits alone let ~100 queries recover attributes never asked — sineadwilliamso · 2026-10-07
- Final logits are far from a minimal information bottleneck, paper warns — sineadwilliamso · 2026-10-07
- Apple paper: top-k logits can leak task-irrelevant info almost as much as residual stream — sineadwilliamso · 2026-10-07
- Study: Top-k Logits Leak as Much Information as Tuned Lens Trajectories, Far More Accessible — sineadwilliamso · 2026-10-07