COLM 2026 paper: yes/no queries leak far more via logits than expected
sineadwilliamso · x · 2026-10-07
The authors will present 'What do your logits know' at COLM 2026, asking a simple but important question: when you query a model with a yes/no question about an image, how much other information leaks through the output? Far more than expected. Thread highlights:
- Residual stream encodes nearly everything, queried or not.
- Even top-2 logits leak: ask 'is there a blue sphere?' and the logits reveal the sphere's material and size, never mentioned in the query.
- Quantified threshold: target attributes are reliably recoverable at k≈10-13; as k grows, background object properties become decodable too.
- API level: just the top-20 logits exposed by many APIs suffice to recover unqueried attributes with as few as 100 queries; comparable to tuned lens trajectories.
- Implications for privacy, fairness, distillation, and interpretability — final logits are far from a minimal information bottleneck.
More from Safety
- OpenAI to watermark ChatGPT text in coming weeks to comply with EU AI Act — paulnovosad · 2026-10-07
- Project Glasswing Reports 135K Verified Vulnerabilities, 9,333 Already Patched — ResultBackground2450 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Pentagon pulls plug on Claude after Anthropic refused to lift limits on surveillance, weapons — mark_k · 2026-10-07
- SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07