Researchers Debate: Why RL Training on Agent Harnesses Generalizes Better
srush_nlp · x · 2026-07-23
A discussion sparked by Sholto Douglas and Rachel Lim's research. Sasha Rush questioned the experiment for merely comparing a raw LLM to one with a Python interpreter, failing to see how it proves the need for token-by-token input. Sholto clarified that the experiment compares direct base LM training against training an LM equipped with a harness, with evaluations deliberately introducing out-of-distribution task lengths or topics. Data shows that RL training on the harness yields better generalization, an effect not seen when arbitrarily applying standard harnesses. He emphasized that the token-by-token input is crucial because the harness shapes the context trajectories for the underlying LM's individual calls.
Related event: LLM Generalization Debate: Intrinsic Model or Harness Contribution?(11 posts)→
More from Research
- Marigold V2 Hits New SOTA in Monocular Depth Estimation with Single-Step Diffusion Transformers — AntonObukhov1 · 2026-09-11
- ZibraAI engineer builds boundless interactive fluid simulation in VRChat, 100K particles run on any PC — Michael_Moroz_ · 2026-09-11
- How AI Agents Turn Experience Into Lasting Gains: A Guide to Recursive Self-Improvement — Roger_M_Taylor · 2026-09-11
- Joshua Gans: ChatGPT 5.2 Pro wrote a full paper in 19 minutes, but quality ideas still matter — joshgans · 2026-09-11
- Four-Color Theorem Gets a Rare New Proof, Revisiting Its Controversial 1970s Computer-Assisted Solution — soumitrashukla9 · 2026-09-11
- The Roadmap of Mathematics for Machine Learning: Linear Algebra, Calculus, Probability — TivadarDanka · 2026-09-11