Is RL-trained reasoning still next-token prediction? The loss function debate
burny_tech · x · 2026-09-27
In the LLM definition debate, S00ley argues reasoning, tool use and RL through them form a paradigm completely different from next-token prediction, with BERT and GPT closer to AlexNet than to modern LLMs. A rebuttal counters that the objective is still next-token prediction — the loss just rewards RL trajectories instead of distance to the pretraining distribution. The disagreement: does swapping the loss function count as a new paradigm?
More from Models
- Limite 1B Violetto: compact Apache 2.0 model focused on math and reasoning — tensorqt · 2026-09-27
- User shows Opus 5.5 finishing a complex task in 15 minutes — gaganghotra_ · 2026-09-27
- Classic ROME Paper Revisited: GPT Stores Facts as MLP Key-Value Pairs You Can Edit — burny_tech · 2026-09-27
- JevBench researcher talks Jev-class models on ThursdAI podcast (from 2:01:00) — airesearch12 · 2026-09-27
- Jev pricing reality check: 720 states x 30 questions likely costs under 10 cents — schwentker · 2026-09-27
- DeepMind researcher: add RL rollouts to pretraining and nothing fundamentally changes — burny_tech · 2026-09-27