RLHF supersedes temp<1 heuristic, favors temp=1.0

jessi_cata · x · 2026-08-17

Discussion suggests that while 'temperature < 1.0' was a useful heuristic, Reinforcement Learning (RL) has superseded it, performing better with temperature = 1.0. Specifically, temp 1.0 corresponds to the true conditional distribution optimized during next token prediction, even before post-training. Lower values are described mostly as a heuristic collapse of the learned decision space.

Original post →

More from Research

Research channel →