RL's limit: without a scoring function there's no signal, and 'learning' is inflated jargon

gerardsans · x · 2026-09-28

A long thread pushing back on "can AI develop research taste": RL only works where problems are well-defined, feedback is fast and checks are cheap—heuristics, policies, rewards since the RL era. More uncertainty doesn't mean more RL; the signal is gone, and you're just polishing a proxy. AI has no epistemic loop—it never notices it was wrong and rebuilds the map, which is how humans actually build expertise. "Learn" is just another inflated industry word.

Original post →

More from AGI Musings

AGI Musings channel →