Tsinghua & Tencent Propose GPS for Reasoning Training

jiqizhixin · x · 2026-07-19

Researchers from Tsinghua University and Tencent propose GPS (Generalizable Predictive Prompt Selection). It uses a lightweight predictive model based on shared training history to first estimate prompt difficulty, then prioritizes moderately difficult and diverse samples to guide the RL post-training of large models.

The paper claims this approach reduces expensive trial-and-error rollouts, outperforming strong baselines on multiple reasoning benchmarks while improving training efficiency, final performance, and test-time speed. The diagram illustrates the complete training pipeline: difficulty prediction, batch selection, generation and reward calculation, RLVR updates, and updates to history and PPM parameters.

Original post →

More from Infra

Infra channel →