Co-occurrence is the mathematical backbone of token-driven LLMs

gerardsans · x · 2026-10-04

Arguing that co-occurrence isn't the full picture but an essential ingredient, the author explains that weights and biases are compressed distributions from gradient descent over a corpus, then breaks down the full chain: embedding creation, inference dynamics (tokenisation, positional encoding, pre-fill, attention/MLP layers, sampling) and the autoregressive loop.

Related event: Debate: Co-occurrence Statistics as the Mathematical Foundation of LLMs(3 posts)→

Original post →

More from Research

Research channel →