After RL training, calling LLMs 'language predictors' is no longer accurate, researcher argues

morqon · x · 2026-09-25

Chris Hayduk argues that since a large share of training compute is no longer spent optimizing next-token prediction, characterizing LLMs as "predicting language" is a mischaracterization — at least once models have undergone substantial RL training.

If the phrase means anything, he says, it should mean "predicting the chain of language that will solve this difficult problem" rather than "predicting the next token" or "being a stochastic parrot." A direct challenge to the popular stochastic-parrot framing of post-RL models.

Original post →

More from Models

Models channel →