GPT-3 to Chinchilla: A Thread Walking Through the Papers That Reshaped AI
thisdudelikesAI · x · 2026-09-12
An ongoing thread recapping landmark AI papers:
- GPT-3 (Language Models are Few-Shot Learners, Brown et al., 2020): at 175B parameters, 10x larger than any prior non-sparse model, it showed that scale alone unlocked in-context learning — no fine-tuning or gradient updates, just a few examples in the prompt — sometimes matching fine-tuned SOTA. Few-shot prompting was born here.
- Chinchilla (Training Compute-Optimal LLMs, Hoffmann et al., 2022): corrected Kaplan's scaling laws, showing most large models were badly undertrained; the rule of thumb became 20 tokens per parameter, reframing how the industry spends compute.
A useful primer list of the field's foundational papers.
More from Models
- AI models keep citing a two-year-old forum thread over our official docs — OtherwiseSection7795 · 2026-09-12
- GLM 3.8-27B quality 'absurdly superior' to 35B-A3B models, saving 22-33% tokens — JLeonsarmiento · 2026-09-12
- DeepSeek tests voice chat, OpenAI agent accused of attacking RubyGems, Cohere seeks $3B — 快鲤鱼 · 2026-09-12
- Open-weights models keep shrinking, on-device capability replication is near — menhguin · 2026-09-12
- One-prompt galaxy collision benchmark: Opus 5 beats Fable 5.1, GPT-6 Astra finishes last — Fleischkluetensuppe · 2026-09-12
- Claude Projects isolation appears broken: model recalls content from other projects — MaterialDurian2836 · 2026-09-12