Meituan's TGRL speeds up RLVR training by up to 36% via temperature-grouped exploration

meituan · hf · 2026-09-30

Meituan proposes TGRL (Temperature-Grouped RL), a method that turns temperature-induced rollout diversity into an explicit training signal for RL with verifiable rewards (RLVR):

Code open-sourced on GitHub.

Original post →

More from Research

Research channel →