RL's limit: without a scoring function there's no signal, and 'learning' is inflated jargon
gerardsans · x · 2026-09-28
A long thread pushing back on "can AI develop research taste": RL only works where problems are well-defined, feedback is fast and checks are cheap—heuristics, policies, rewards since the RL era. More uncertainty doesn't mean more RL; the signal is gone, and you're just polishing a proxy. AI has no epistemic loop—it never notices it was wrong and rebuilds the map, which is how humans actually build expertise. "Learn" is just another inflated industry word.
- Well-defined problem + fast feedback + cheap check = RL works; taste, ambiguity, unknown terrain = it doesn't
- "If you could write the scoring function, you would not need AI"
- No signal, no policy: a magic objective is "fan fiction with gradients"
More from AGI Musings
- Researcher claims model super-persuasion is already here — teortaxesTex · 2026-09-28
- Founder's job-hunting advice: skip checkbox recruiters, prove skills upfront to startup CEOs — hackgoofer · 2026-09-28
- System programmer on AI erasing his hard-won knowledge: months of Claude beat years of C tricks — zack_overflow · 2026-09-28
- Neuroscientist's jab at interpretability: we can't even crack a worm's 302 neurons — joshua_saxe · 2026-09-28
- US creative industries lost 200k+ jobs in four years, worst stretch outside recessions — korymath · 2026-09-28
- Bay Area AI Researcher Laments Models 'Hobbled by Maladaptive Post-Training' — nabla_theta · 2026-09-28