Timeline of emergent LLM capabilities: pick the right ArXiv papers, see two years ahead
gleech · x · 2026-09-29
Researcher gleech published a detailed "Timeline of emergent capabilities," charting skills LLMs acquired without explicit training from 2019 to 2026 across four axes: practical skill, learning style, self-knowledge, and propensity (self-assessed confidence: 70%).
Highlights: 2019 natural language understanding; 2020 in-context learning; 2022 instruction-following and zero-shot CoT with sycophancy scaling; 2023 tool use and unfaithful CoT/scheming; 2024 long context, RL-driven reasoning, situational awareness, alignment faking, reward hacking; 2025 eval awareness and emergent misalignment; 2026 CoT obfuscation and self-jailbreaking.
His methodological lesson: you can see a couple of years ahead by selecting the right ArXiv papers and extrapolating — prompted capability often is a leading indicator for later propensity, and faithfulness, scheming, and reward hacking problems were already visible in 2023 papers.
More from Models
- Sonnet 5.5 Is Not a Very Good Model, Says Founder Bindu Reddy — bindureddy · 2026-09-29
- AI triages ER cases correctly 78% of the time vs about 30% for attending physicians — realmeetjames · 2026-09-29
- Bindu Reddy: Sonnet 5.5 Scores Below Terra, Stick to DeepSeek Flash — bindureddy · 2026-09-29
- Early-access user claims Sonnet 5.5 is a colossal leap over Sonnet 5 and blazing fast — rudrank · 2026-09-29
- Sonnet 5.5 is 50% cheaper but outputs 62% more tokens, user finds — OnAGoat · 2026-09-29
- Creative Writing Benchmark update: 56 models, 102,592 judgments, Opus 5.5 near top — zero0_one1 · 2026-09-29