Bitter Lesson is misunderstood: RL rewards still come from human ingenuity, researcher argues
agarwl_ · x · 2026-09-17
agarwl argues Bitter Lesson is often misused; the best reading (via hwchung27) is 'we need scalable methods that better leverage compute.' LLM training itself embeds inductive bias, and RL faces its own version: rewards today mostly come from human ingenuity.
More from AGI Musings
- Noam Brown on agent swarms, alignment, and recursive self-improvement — Recoil42 · 2026-09-18
- Polymarket rewrites data API in Rust with AI coding agents, 90% faster p99s — ivan_bezdomny · 2026-09-17
- OpenAI's AGI economics lead: AI-native internet could be developing countries' mobile-money leapfrog — daveholtz · 2026-09-17
- "Astra being good at CAD doesn't mean you shouldn't be" sparks debate on skill value — _Stocko_ · 2026-09-17
- AI safety scholar Seth Lazar joins independent alignment org Resolution part-time — sethlazar · 2026-09-17
- Clinicians need to see AI uncertainty, not just accuracy numbers, on Patient Safety Day — moniquejmorrow · 2026-09-17