Final logits are far from a minimal information bottleneck, paper warns
sineadwilliamso · x · 2026-10-07
The thread stresses direct implications for privacy, fairness, distillation, and interpretability: a model's final logits are far from a minimal information bottleneck. Combined with earlier findings, just the top-20 logits exposed by APIs suffice to recover unqueried attributes with as few as 100 queries.
More from Safety
- OpenAI to watermark ChatGPT text in coming weeks to comply with EU AI Act — paulnovosad · 2026-10-07
- Project Glasswing Reports 135K Verified Vulnerabilities, 9,333 Already Patched — ResultBackground2450 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Pentagon pulls plug on Claude after Anthropic refused to lift limits on surveillance, weapons — mark_k · 2026-10-07
- SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07