Zhejiang University and Tencent’s STEER targets entropy collapse in RL training

jiqizhixin · x · 2026-07-24

Zhejiang University and Tencent propose STEER to fix entropy collapse in RL for LLMs

Researchers from Zhejiang University and Tencent introduce STEER, a method aimed at a common failure mode in reasoning training: entropy collapse.

Instead of using a heuristic to blindly adjust entropy, STEER estimates how entropy changes and then reweights tokens adaptively. The authors claim this targets a key flaw in RL for LLMs more directly than prior interventions.

According to the post, the method outperforms state-of-the-art baselines on:

The paper is titled Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective and the code is open-sourced.

Original post →

More from Research

Research channel →