Researchers Debate: Why RL Training on Agent Harnesses Generalizes Better

srush_nlp · x · 2026-07-23

A discussion sparked by Sholto Douglas and Rachel Lim's research. Sasha Rush questioned the experiment for merely comparing a raw LLM to one with a Python interpreter, failing to see how it proves the need for token-by-token input. Sholto clarified that the experiment compares direct base LM training against training an LM equipped with a harness, with evaluations deliberately introducing out-of-distribution task lengths or topics. Data shows that RL training on the harness yields better generalization, an effect not seen when arbitrarily applying standard harnesses. He emphasized that the token-by-token input is crucial because the harness shapes the context trajectories for the underlying LM's individual calls.

Related event: LLM Generalization Debate: Intrinsic Ability or Harness Illusion(10 posts)→

Original post →

More from Research

Research channel →