Two Years After o1, Reasoning Tokens Are Table Stakes for All Frontier Models
CurieuxExplorer · x · 2026-09-26
Artificial Analysis marks two years since o1-preview, the first reasoning model that made models pause and think before answering—now standard across all frontier models.
- v1 of its Intelligence Index used just four single-turn exam-style evals: MMLU, GPQA, MATH, HumanEval
- Today's v4.3 incorporates 10 difficult evaluations covering long-horizon agentic tasks, hard coding, and knowledge work
- A snapshot of how the industry's eval focus shifted from Q&A to agentic capability
Related event: Two Years After o1-preview, Reasoning Is Standard Across Frontier Models(3 posts)→
More from Models
- Opus 5.5 does worse and costs more at Max thinking — Medium beats Max on hard benchmark — davidyin44 · 2026-09-26
- LibertAI ships open-weight Deem 9B on Qwen3.5, trailing Jev 68.9% vs 74.1% — Pokenhagen · 2026-09-26
- Peter Steiberger says he codes with Codex plus an OC harness, not Claude — steipete · 2026-09-26
- Why Anthropic lags OpenAI on math: compute constraints and 400k GPUs coming online — haider1 · 2026-09-26
- Hytale WorldGen V2 face-off: Claude Opus 4.6 vs Opus 5.5 compared — Angaisb_ · 2026-09-26
- OpenAI's rumored persistent agent "o" may tie to old "rebranding to O" report — Dullydude · 2026-09-26