RELER trains embedding models with RL directly in embedding space, beating InfoNCE
_reachsumit · x · 2026-10-07
RELER is a reinforcement learning framework that lets existing embedding models optimize retrieval metrics and task rewards directly in embedding space. It samples unit-length embeddings from vMF distributions centered on encoder outputs, scores retrieval outcomes as rewards, and updates via REINFORCE with RLOO. A conditional-mean projection (CMP) reduces policy-gradient noise in high-dimensional exploration. On the reasoning-intensive BRIGHT benchmark, RELER consistently beats InfoNCE and LambdaLoss when post-training BGE-M3 and Qwen3-Embedding, and improves downstream RAG utility.
More from Research
- Andrew Davison: robots need object-based SLAM, not scan-then-fit reconstructions — AjdDavison · 2026-10-07
- CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining — Shangye Song · 2026-10-07
- Training-Free Accent Analogy Guidance Boosts Speaker Similarity in Cross-Lingual Voice Cloning — Yoomee Cho · 2026-10-07
- Source Attribution of Synthetic Data Hits 98.7% Accuracy but Falls to 29% After Style Rewriting — Joss Armstrong · 2026-10-07
- Physicist finds fractal patterns (D 1.3-1.5) cut stress response by up to 60% — aakashgupta · 2026-10-07
- AI has now cracked at least 10 open math problems each worthy of a Fields Medal — luismbat · 2026-10-07