RELER trains embedding models with RL directly in embedding space, beating InfoNCE

_reachsumit · x · 2026-10-07

RELER is a reinforcement learning framework that lets existing embedding models optimize retrieval metrics and task rewards directly in embedding space. It samples unit-length embeddings from vMF distributions centered on encoder outputs, scores retrieval outcomes as rewards, and updates via REINFORCE with RLOO. A conditional-mean projection (CMP) reduces policy-gradient noise in high-dimensional exploration. On the reasoning-intensive BRIGHT benchmark, RELER consistently beats InfoNCE and LambdaLoss when post-training BGE-M3 and Qwen3-Embedding, and improves downstream RAG utility.

Original post →

More from Research

Research channel →