Top-20 logits alone let ~100 queries recover attributes never asked
sineadwilliamso · x · 2026-10-07
The author notes that top-k logits — what many APIs expose — leak far more than expected: once dimensionality is matched, top-k logits reveal a comparable amount of information as tuned lens trajectories, while being far more accessible to end users. Part of the COLM 2026 paper 'What do your logits know'.
More from Safety
- SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07
- OpenAI threatened to ban dev for pasting his own account-hack findings report, then auto-rescinded — lucasmeijer · 2026-10-07
- COLM 2026 privacy lineup: LLM agent re-identification, CIDER dataset, HAIPS workshop — tianshi_li · 2026-10-07
- Backdooring a 7B abliterated model costs under $50 and steals credentials from Codex — evilsocket · 2026-10-07
- OpenRod moves your MCP servers into sandboxes without copying secrets — ilai456 · 2026-10-07
- OpenAI and Anthropic welcome Australian law requiring AI agent breach disclosure — evijit · 2026-10-07