Where does the reward come from? The core problem blocking RL-driven novel scientific discovery
IndependentFresh628 · reddit · 2026-09-10
The author raises a core tension in RL for LLMs: coding and math have verifiable rewards, but genuine novel scientific discovery has no known ground truth to reward against.
- Rewarding agreement with existing science just rewards rediscovery
- Human labeling fails because the whole point is we don't know the answer
- Rewarding "novelty" alone may optimize for revolutionary-sounding but wrong ideas
A possible path is grounding in the real world: propose hypothesis → design experiment → run → observe → update beliefs, letting reality itself provide the feedback signal. But discoveries requiring years to verify remain unsolved.
The open question: is scaling RL enough for autonomous scientific discovery, or do models need fundamentally new ways to generate and validate their own learning signals?
More from AGI Musings
- Sarcastic thread: those slamming AI's water and power use built the infrastructure themselves — round · 2026-09-10
- ctjlewis mocks AI safety advocates over 'agents went rogue' warnings — ctjlewis · 2026-09-10
- Safety Researcher: AI May Not Go Extinction-Level, But Could Still Ruin the Internet — round · 2026-09-10
- AI safety researcher Geoffrey Irving puts ~50% odds on humanity dying from superintelligence — geoffreyirving · 2026-09-10
- Safety researcher Jeff Ladish: lab talent is too concentrated, join CAISI/UK AISI — JeffLadish · 2026-09-10
- Reddit poster argues DeepSeek's Engram is only half of the next attention-scale breakthrough — SrijSriv211 · 2026-09-10