USTC and Alibaba Propose PARPO: Personalized Agentic RL That Optimizes Preference, Not Just Correctness

量子位 · wechat · 2026-09-13

Researchers from USTC and Alibaba propose PARPO (Personalized Anchor Reward-Decoupled Policy Optimization), formally bringing personalization into the training objective of Agentic RL.

Problem

Method

Results

Limitations: limited human blind-test scale; the interest-conformity decoupling is operational, not strict causal separation.

Original post →

More from Research

Research channel →