How LLMs actually work: embeddings, inference dynamics and the autoregressive loop, explained

gerardsans · x · 2026-10-04

In a reply thread, the author breaks down "how LLMs work" into three parts: embeddings creation (token configuration shaped by training regimes), inference dynamics (tokenization, positional encoding, pre-fill, layer-to-layer attention and MLP blocks leading to token sampling), and the autoregressive loop. A linked article series provides the full explanation — a solid primer on the Transformer inference pipeline.

Related event: Debate: Co-occurrence Statistics as the Mathematical Foundation of LLMs(3 posts)→

Original post →

More from Research

Research channel →