Microsoft Paper: Retrieval-Grounded Credit Assignment Fixes Sparse Rewards in Generative Recommenders

_reachsumit · x · 2026-09-25

A Microsoft team published an arXiv paper addressing the credit-assignment gap in reasoning-enhanced generative recommenders. Training with exact-match Semantic ID rewards under GRPO is sparse: groups where all rollouts miss the target yield zero learning signal, and rollouts sharing the same SID reward get identical advantages regardless of their reasoning traces. The fix structures each trace into a history summary, interest hypotheses, and a final SID, then executes every hypothesis as a catalog query with a frozen retriever—rewarding rollouts where any query hits the target in top-K and localizing credit to individual hypotheses for span-level credit assignment.

Original post →

More from Research

Research channel →