'Next-token predictor' is a contentless way to describe LLMs

Aaroth · x · 2026-09-25

Aaroth points out an annoying part of the Twitter/Bluesky debate over whether LLMs are "only" next-token predictors: the phrase is contentless, since all mappings from inputs to output strings can be factored into a sequence of next-token distributions. A follow-up adds this is a boring syntactic fact about probability distributions, not a claim about how brains work.

Original post →

More from AGI Musings

AGI Musings channel →