vec2vec translates any two text-embedding spaces without paired data, hits 0.96 cosine
maier_ak · x · 2026-09-10
Cornell's vec2vec learns a translation between any two text-embedding spaces without paired sentences, the original encoders, or a dictionary, using adversarial training, reconstruction, cycle-consistency and distance-preserving losses. On the hardest cross-backbone pairs it reaches up to 0.96 cosine similarity, 100% top-1 accuracy and mean rank 1; on out-of-distribution tweets and clinical notes it still exceeds 0.73 cosine, far above random guessing.
More from Research
- Meta's Boxer at ECCV: closing the 3D ground-truth gap with 2D scaling — ducha_aiki · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Hugging Face shows how to train on 1M-token sequences on a single 8-GPU node — QGallouedec · 2026-09-10
- SyncWorld turns world models into zero-shot robot simulators via visual calibration — Yuncong Yang · 2026-09-10
- Feng Yao wins ECVA PhD Award at ECCV 2026 for 3D humans + language thesis — Michael_J_Black · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10