Pretraining is just RL with single-token rollouts and full-information feedback

burny_tech · x · 2026-09-27

willcb offers a technical observation: every LLM starts as a random-init model, and pretraining is formally equivalent to RL with single-token rollouts and full-information (rather than bandit) feedback, unifying next-token prediction with the RL framework.

Original post →

More from Research

Research channel →