HF Engineer Tests: Increasing Rollout Token Budget Improves Model Win Rate

mervenoyann · x · 2026-08-13

Hugging Face engineer Merve Noyan shared her experiments training Nemotron3.5 on the game of Wordle using OpenEnv and TRL.

She observed that increasing the maximum token limit per rollout from 2048 to 4096 improved the model's base win rate from 52% to 68%. She originally intended to train the model to be more token-efficient, but the model kept hitting the completion length limit during rollouts, causing the attempt to fail and only increasing side rewards.

Original post →

More from Research

Research channel →