Where does the reward come from? The core problem blocking RL-driven novel scientific discovery

IndependentFresh628 · reddit · 2026-09-10

The author raises a core tension in RL for LLMs: coding and math have verifiable rewards, but genuine novel scientific discovery has no known ground truth to reward against.

A possible path is grounding in the real world: propose hypothesis → design experiment → run → observe → update beliefs, letting reality itself provide the feedback signal. But discoveries requiring years to verify remain unsolved.

The open question: is scaling RL enough for autonomous scientific discovery, or do models need fundamentally new ways to generate and validate their own learning signals?

Original post →

More from AGI Musings

AGI Musings channel →