Train on Frontier Papers or Build RL Envs? An Insider Debate on Math Model Training

ctjlewis · x · 2026-09-10

In a thread questioning whether labs should train directly on frontier mathematicians' results, ctjlewis argues the better path: give agents access to frontier research (from OpenAI and Anthropic) as reference, and test training on the ability to reproduce recent cutting-edge papers — rather than feeding solutions verbatim. His point: building environments and forcing first-principles reproduction internalizes capability, while direct training on answers is just verbatim theft of the solution.

Related event: Debate: Feed Frontier Papers to Training or Reproduce from Scratch(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →