Tsinghua & Tencent Propose GPS for Reasoning Training
jiqizhixin · x · 2026-07-19
Researchers from Tsinghua University and Tencent propose GPS (Generalizable Predictive Prompt Selection). It uses a lightweight predictive model based on shared training history to first estimate prompt difficulty, then prioritizes moderately difficult and diverse samples to guide the RL post-training of large models.
The paper claims this approach reduces expensive trial-and-error rollouts, outperforming strong baselines on multiple reasoning benchmarks while improving training efficiency, final performance, and test-time speed. The diagram illustrates the complete training pipeline: difficulty prediction, batch selection, generation and reward calculation, RLVR updates, and updates to history and PPM parameters.
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21