Beyond Cross-Entropy: Pre-training Language Models with Pure Reinforcement Learning

tokenbender · x · 2026-08-01

Traditional language model pre-training relies on cross-entropy, rigidly teaching that only a single token is the correct continuation. The article introduces avatarl, a new paradigm proposing to replace this with a reinforcement learning (RL) framework during pre-training.

Related event: Explorative Modeling Introduces Third Pretraining Axis, Questioned as avataRL Reinvention(8 posts)→

Original post →

More from Research

Research channel →