Alibaba's SSR-GRPO: RL for E-commerce Dense Retrieval

_reachsumit · x · 2026-08-21

Alibaba presents SSR-GRPO, a reinforcement learning method for e-commerce search. To address bias from using LLMs as reward models, SSR-GRPO leverages Semantic IDs and dense vectors for unbiased relevance scoring. It also mines hard negatives to mask noisy samples, improving the handling of complex and implicit semantics.

Original post →

More from Research

Research channel →