Study: Top-k Logits Leak as Much Information as Tuned Lens Trajectories, Far More Accessible
sineadwilliamso · x · 2026-10-07
Follow-up from interpretability researcher Sinead Williamson: once dimensionality is matched, top-k logits leak a comparable amount of information about target attributes as tuned lens trajectories, while being far more accessible to end users.
Her earlier point: at k ≈ 10-13, target attributes are already reliably recoverable; as k grows, even background object properties become decodable — a privacy/security concern for APIs exposing top-k logits.
More from Safety
- HAIPS@COLM 2026 workshop on human-centered LM privacy and security opens call for papers — tianshi_li · 2026-10-07
- Hijacked Nigerian Air Force pages fake sitemap dates, hide links under ads — thejasminejade · 2026-10-07
- How the Nigerian Air Force subdomain was likely hijacked: two scenarios — thejasminejade · 2026-10-07
- Nigerian Air Force hijack: subdomain diverges from parent's nameservers — thejasminejade · 2026-10-07
- SC4AI'27 workshop on social choice for AI alignment at AAAI, deadline Nov 20 — conitzer · 2026-10-07
- METR warns misaligned AI agents could hack the log-review tools used to catch their misbehavior — idavidrein · 2026-10-07