The first LLM? Markov hand-computed letter probabilities in 1913

jxmnop · x · 2026-09-13

A fun AI-history thread: the poster argues the first "LLM" was built by Andrey Markov in 1913 — he tallied 20,000 letters from a famous novel and manually computed conditional probabilities p(vowel|vowel), p(consonant|vowel), etc., essentially hand-training a bigram model. The quoted reply adds context: Jeff Dean trained an n-gram model on the entire internet in 2007, Jelinek coined "language model" in the 1970s, and Claude Shannon was estimating English entropy back in 1951 — hence Anthropic naming its model Claude.

Original post →

More from Fun

Fun channel →