Kuaishou's SARA Scales Natural-Language Preference Rationales From 240M Users for Recommendations
_reachsumit · x · 2026-09-17
A Kuaishou arXiv paper introduces SARA, an industrial framework using articulated user rationales (AURs) — users' natural-language explanations of why they like or dislike content — as a new recommendation signal, beyond implicit behaviors like clicks and watch time that reveal what users do but not why.
AURs are sparse, noisy, and low-coverage, so SARA:
- Builds a data engine eliciting and curating AURs from 240M Kuaishou Live users into SARA-HQ, a quality-controlled, author-centric dataset.
- Aligns a general MLLM into SARA-7B via large-scale SFT and Quality-Refining DPO, extending rationale generation from 86,564 covered authors to the full 10M-author space.
- Feeds generated positive/negative rationales into production ranking through SARA-Ranker with rationale-aware interaction modeling and rejection-memory modeling.
More from Research
- Rumors: Jev model tackled bounded prime gaps, FLT formalization and Navier-Stokes in a week — f_charton · 2026-09-20
- Regression diagnostics: VIF ≥ 10 means collinearity is inflating your standard errors — mdancho84 · 2026-09-20
- 6 Regression Diagnostic Checks Every Data Scientist Should Run — mdancho84 · 2026-09-20
- EMNLP paper: LM rewrites inflate certainty in up to 75% of outputs, 1.5-2x bias toward stronger claims — sameer_ · 2026-09-20
- Pachter Lists More Flaws in Nucleus Ads: Implausible R² Matches, Mismatched Cohorts — lpachter · 2026-09-20
- Biologist Debunks Nucleus Claim That Embryo Screening Boosts IQ by 10.8 Points — lpachter · 2026-09-20