Doing RL with LLMs is like rediscovering human nature from first principles

burny_tech · x · 2026-08-12

A developer shared insights from working on Reinforcement Learning (RL) with LLMs, noting that the process feels like rediscovering human nature and social constructs from first principles.

They observed that in the training loop, providing words of encouragement—similar to human reinforcement—actually matters and yields positive results.

Original post →

More from Fun

Fun channel →