HF Engineer Tests: Increasing Rollout Token Budget Improves Model Win Rate
mervenoyann · x · 2026-08-13
Hugging Face engineer Merve Noyan shared her experiments training Nemotron3.5 on the game of Wordle using OpenEnv and TRL.
She observed that increasing the maximum token limit per rollout from 2048 to 4096 improved the model's base win rate from 52% to 68%. She originally intended to train the model to be more token-efficient, but the model kept hitting the completion length limit during rollouts, causing the attempt to fail and only increasing side rewards.
More from Research
- Embodied AI Breakthrough: SONIC System for Robot Motion Tracking Published in Science Robotics — zhengyiluo · 2026-08-13
- Study Shows Diminishing Returns to LLM Intelligence, Challenging Frontier Model Premiums — soumitrashukla9 · 2026-08-13
- Study: CLAUDE.md files grow unbounded; comments cut 99.3% excess instructions — omarsar0 · 2026-08-13
- NVIDIA Launches AI-Aided Engineering Group to Accelerate Physical System Design — JeanKossaifi · 2026-08-13
- NeurIPS 2026 MATH-AI Workshop Calls for Papers on Agentic AI and Math — KaiyuYang4 · 2026-08-13
- Automating AI Research: A Retrospective on Breaking World Records — tensorqt · 2026-08-13