Two years since o1-preview: reasoning tokens went from novelty to frontier standard

ArtificialAnlys · x · 2026-09-26

Artificial Analysis marks two years since its o1-preview report: o1 was the first model to substantially push the intelligence frontier since GPT-4 and the first reasoning model; today every frontier model uses reasoning tokens to think before answering. Evaluation methodology evolved in parallel — Intelligence Index v1 used four single-turn exam-style benchmarks (MMLU, GPQA, MATH, HumanEval), while v4.3 now runs 10 harder evals covering long-horizon agentic tasks, challenging coding, and knowledge work. The original report also flagged o1's speed and cost trade-offs as making it unsuitable for most production use at the time.

Related event: Two Years After o1-preview, Reasoning Is Now Standard at the Frontier(2 posts)→

Original post →

More from Models

Models channel →