vec2vec translates any text embedding spaces without paired data, exposing stolen vector DBs to privacy attacks

A Cornell team has released the paper "Reading Between the Lines: Translating Text Embeddings Across Model Borders," introducing vec2vec, a method that learns to translate between any two text embedding spaces without paired sentences, the original encoders, or a dictionary. Relying only on adversarial training, reconstruction, cycle-consistency, and distance-preserving losses, it exploits the universal geometric structure of embeddings to complete the mapping. The result shows that embedding spaces from different models share exploitable commonalities, posing security risks for vector databases.

Confirmed

Why it matters

2026-09-10 ~ 2026-09-10 · 5 related posts

Primary sources