Stolen vector databases weaponized: embedding translation recovers sensitive data at 80% accuracy

maier_ak · x · 2026-09-10

A deep-dive on Cornell's "Harnessing the Universal Geometry of Embeddings" (NeurIPS 2025): vec2vec learns to translate embeddings between model spaces using geometry alone—no paired examples, encoders or dictionary—supporting the Platonic Representation Hypothesis. Security-wise, a stolen vector database can be weaponized: translating embeddings into a known space enables zero-shot attribute inference and inversion, recovering sensitive facts with up to 80% accuracy on Enron emails without any access to the original encoder.

Related event: vec2vec translates any text embedding spaces without paired data, exposing stolen vector DBs to privacy attacks(5 posts)→

Original post →

More from Safety

Safety channel →