Reading List: Evolution of LLM Architectures
craigsdennis · x · 2026-07-15
This post curates a list of resources to understand the evolution of LLM architectures, covering the trajectory from early GPT-style models to subsequent key concepts.
Highlights include:
- An article reviewing 5 years of progress in GPT-style models, stopping before 2023 to offer a clearer view of early model details.
- A companion post praised for its visualizations and code examples.
- A dense, math-heavy survey detailing how key LLM concepts gradually took shape.
- Another article that serves as both a beginner-friendly intro to "what is a token" and a overview of early representative LLMs.
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- Structural ensembles, not single predictions, drive robust TCR:pMHC generalization — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22
- LLM leaderboards are now often measuring the harness too, Gary Marcus warns — GaryMarcus · 2026-07-22
- New paper defines self-state attacks, showing OS defenses leave four agent-memory cases indistinguishable — Justgototheeffinmoon · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22