Four years of AI progress across seven capabilities, adjusted for benchmark changes
epheva · reddit · 2026-10-04
A Reddit gallery charts four years of AI progress across seven capability dimensions, with adjustments for changes in the benchmarks themselves — aiming for a fairer year-over-year comparison than raw leaderboard scores. The key data points are in the attached images.
More from Models
- Grok 4.7 tops Artificial Analysis' Cyber Index, beating Claude Opus 5.5 and ChatGPT 6 Astra — XFreeze · 2026-10-04
- One-Shot Brutalist City Builder: Fable 5.1 Shows a Year of Rapid Capability Gains — Afinetheorem · 2026-10-04
- First confirmed out-of-app Grok jailbreak claimed, security community takes notice — kleffew94 · 2026-10-04
- 3.6B TwIL-LM3-Pro runs locally on 4GB VRAM, claims 35% lead over VibeThinker-3B — glenbeer · 2026-10-04
- Elon Musk confirms Grok 4.7 'writes great code' in one-word reply — elonmusk · 2026-10-04
- Kolibri team on data selection: blending multiple quality signals instead of one score — josh_wills · 2026-10-04