Study: Top-k Logits Leak as Much Information as Tuned Lens Trajectories, Far More Accessible

sineadwilliamso · x · 2026-10-07

Follow-up from interpretability researcher Sinead Williamson: once dimensionality is matched, top-k logits leak a comparable amount of information about target attributes as tuned lens trajectories, while being far more accessible to end users.

Her earlier point: at k ≈ 10-13, target attributes are already reliably recoverable; as k grows, even background object properties become decodable — a privacy/security concern for APIs exposing top-k logits.

Related event: Study: VLM top-k logits leak far more task-irrelevant image information than expected(8 posts)→

Original post →

More from Safety

Safety channel →