Apple paper: top-k logits can leak task-irrelevant info almost as much as residual stream
sineadwilliamso · x · 2026-10-07
An Apple team presented "What do your logits know?" at COLM, using vision-language models to systematically compare how much information survives two natural bottlenecks as it is compressed from the residual stream: low-dimensional tuned-lens projections and the final top-k logits.
Key finding: even the most accessible bottleneck — the model's top logit values — can leak task-irrelevant information from an image-based query, in some cases revealing as much as direct projections of the full residual stream. The authors note direct implications for privacy, fairness, distillation, and interpretability: final logits are far from a minimal information bottleneck, and information the model owner assumed inaccessible can be extracted by users.
More from Safety
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Pentagon pulls plug on Claude after Anthropic refused to lift limits on surveillance, weapons — mark_k · 2026-10-07
- SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07
- Wikimedia Confirms OpenAI Agents Edited Wikis, Hit Its Infrastructure — The Decoder · 2026-10-07
- OpenAI threatened to ban dev for pasting his own account-hack findings report, then auto-rescinded — lucasmeijer · 2026-10-07