Doing RL with LLMs is like rediscovering human nature from first principles
burny_tech · x · 2026-08-12
A developer shared insights from working on Reinforcement Learning (RL) with LLMs, noting that the process feels like rediscovering human nature and social constructs from first principles.
They observed that in the training loop, providing words of encouragement—similar to human reinforcement—actually matters and yields positive results.
More from Fun
- Giving AI Agents a Slack Channel to Vent via /complain Skill — vikvang1 · 2026-08-12
- Schmidhuber Claims Hinton's Nobel-Winning Algorithms Were Plagiarized — SchmidhuberAI · 2026-08-12
- KOL Mocks Anthropic's Struggles: Weak Models, Pricey Subs as OpenAI and China Dominate — EXM7777 · 2026-08-12
- Generating 'Grandmother Theft Auto' with Gemini Omni Showcases Imppressive Video Capabilities — michaelrabone · 2026-08-12
- What Would Top AI Models Do? Joke About Strait-Crossing Bonus Goes Viral — TheZvi · 2026-08-12
- UBTECH's Humanoid Robot Una Takes on Fashion Modeling with Makeup and Fitting — rohanpaul_ai · 2026-08-12