Alibaba's SSR-GRPO: RL for E-commerce Dense Retrieval
_reachsumit · x · 2026-08-21
Alibaba presents SSR-GRPO, a reinforcement learning method for e-commerce search. To address bias from using LLMs as reward models, SSR-GRPO leverages Semantic IDs and dense vectors for unbiased relevance scoring. It also mines hard negatives to mask noisy samples, improving the handling of complex and implicit semantics.
More from Research
- ShikharMurty: Pass@k gains may just sharpen distribution — ShikharMurty · 2026-08-21
- FasterFASTA tool achieves 1.69 GB/s with multithreaded BGZ decoding — viglovikov · 2026-08-21
- Princeton Paper: Legal Search Benchmarks Fail in Practice, New Dataset Released — burkov · 2026-08-21
- MIT Paper Finds Deleting Artist Data Doesn't Stop AI Recreating Images — technollama · 2026-08-21
- New journal to adopt GEB board, questioning value of legacy publishers — Afinetheorem · 2026-08-21
- HarnessEval-W: An Agentic Benchmark for World Models Evaluation — 青稞AI · 2026-08-21