Meituan's TGRL speeds up RLVR training by up to 36% via temperature-grouped exploration
meituan · hf · 2026-09-30
Meituan proposes TGRL (Temperature-Grouped RL), a method that turns temperature-induced rollout diversity into an explicit training signal for RL with verifiable rewards (RLVR):
- For each prompt, the rollout group is split into low- and high-temperature subsets; exploration gain is estimated from their reward contrast, then allocated as token-level credit using JS divergence between the temperature-scaled next-token distributions.
- Results: equivalent accuracy up to 36% faster than strong RLVR baselines without expanding rollout budget. Across 11 benchmarks: +1.6% six-benchmark math average at 32B, +196.7 CodeForces rating, +4.4% LiveCodeBench Pass@16, and +6.3%/+4.9% ALFWorld/WebShop success rates.
Code open-sourced on GitHub.
More from Research
- USC's CLAM learns robot policies from unlabeled videos, 2-3x success over baselines — ebiyik_ · 2026-09-30
- Editable artifacts may beat screenshots for testing agents' visual understanding — OliviaYii · 2026-09-30
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- VLANeXt Family: 500+ controlled experiments distill 12 practical recipes for VLA models — ccloy · 2026-09-30
- RL with Confidence Margin: COLM 2026 paper makes step-by-step confidence track reasoning correctness — EliasEskin · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30