Is RL-trained reasoning still next-token prediction? The loss function debate

burny_tech · x · 2026-09-27

In the LLM definition debate, S00ley argues reasoning, tool use and RL through them form a paradigm completely different from next-token prediction, with BERT and GPT closer to AlexNet than to modern LLMs. A rebuttal counters that the objective is still next-token prediction — the loss just rewards RL trajectories instead of distance to the pretraining distribution. The disagreement: does swapping the loss function count as a new paradigm?

Original post →

More from Models

Models channel →