Meta Introduces DASO to Optimize Training Signals for Generative Recommenders

_reachsumit · x · 2026-08-24

Meta presents Difficulty-Aware Semantic-ID Optimization (DASO), a post-training method addressing the failure mode of vanilla GRPO in tree-structured generative recommendation tasks. Traditional Semantic-ID based recommendation treats retrieval as autoregressive generation over hierarchical item identifiers, often using SFT followed by GRPO. However, diagnostics reveal that many prompts miss the exact target in the top candidates, leading to weak reward signals. DASO profiles rollout groups by prefix-match depth to locate bottlenecks, reallocates a portion of the group to prefix-guided completions, and utilizes a SID-prefix reward alongside an auxiliary SFT anchor. The method shows improved performance on public benchmarks.

Original post →

More from Research

Research channel →