Data-dependent length penalties and verifier instability discussed by RL trainers
stochasticchasm · x · 2026-09-22
A technical discussion on length penalty design in RL training:
- The author favors group-relative length penalty formulations, since length penalties should be data-dependent—some tasks simply require more tokens than others.
- He is torn between the group-relative approach and calibrating against a reference trace as in k3.
- Notes that verifier problems can cause far more training instability than expected, with compelling graphs to back it up.
Relevant to anyone following RL training practice for reasoning models.
More from Research
- Berkeley and DeepMind Propose Instruct-to-Act: Making World-Model Controllers Language-Instructable — berkeley_ai · 2026-09-22
- Why OpenAI bets on math: it's the most verifiable domain for reinforcement learning — burny_tech · 2026-09-22
- Must-read papers of the week: recursive self-improvement, world models, KV cache compression — TheTuringPost · 2026-09-22
- Meta's A-MLE agent automates ML experimentation for ads ranking, cutting error 2.56% — rohanpaul_ai · 2026-09-22
- Blind RSA apps like Privacy Pass face real-world threat model from scaled oracle queries — matthew_d_green · 2026-09-22
- Harvard/MIT paper FINSKILLOPS makes financial AI self-improve via regression-tested skills — rohanpaul_ai · 2026-09-22