vec2vec translates any two text-embedding spaces without paired data, hits 0.96 cosine

maier_ak · x · 2026-09-10

Cornell's vec2vec learns a translation between any two text-embedding spaces without paired sentences, the original encoders, or a dictionary, using adversarial training, reconstruction, cycle-consistency and distance-preserving losses. On the hardest cross-backbone pairs it reaches up to 0.96 cosine similarity, 100% top-1 accuracy and mean rank 1; on out-of-distribution tweets and clinical notes it still exceeds 0.73 cosine, far above random guessing.

Related event: vec2vec translates any text embedding spaces without paired data, exposing stolen vector DBs to privacy attacks(5 posts)→

Original post →

More from Research

Research channel →