Bitter Lesson is misunderstood: RL rewards still come from human ingenuity, researcher argues

agarwl_ · x · 2026-09-17

agarwl argues Bitter Lesson is often misused; the best reading (via hwchung27) is 'we need scalable methods that better leverage compute.' LLM training itself embeds inductive bias, and RL faces its own version: rewards today mostly come from human ingenuity.

Original post →

More from AGI Musings

AGI Musings channel →