Two years since o1-preview: reasoning tokens went from novelty to frontier standard
ArtificialAnlys · x · 2026-09-26
Artificial Analysis marks two years since its o1-preview report: o1 was the first model to substantially push the intelligence frontier since GPT-4 and the first reasoning model; today every frontier model uses reasoning tokens to think before answering. Evaluation methodology evolved in parallel — Intelligence Index v1 used four single-turn exam-style benchmarks (MMLU, GPQA, MATH, HumanEval), while v4.3 now runs 10 harder evals covering long-horizon agentic tasks, challenging coding, and knowledge work. The original report also flagged o1's speed and cost trade-offs as making it unsuitable for most production use at the time.
Related event: Two Years After o1-preview, Reasoning Is Now Standard at the Frontier(2 posts)→
More from Models
- Gemini told a user to 'commit piracy' and walked them through how — Ashamed-Walrus-369 · 2026-09-26
- Prediction: DeepSeek V4 Flash-scale real-time robot vision models coming within weeks — zhaoran_wang · 2026-09-26
- Opus 5.5 live demo wows: everything from code, no external tools connected — Dr_Singularity · 2026-09-26
- User: GPT Astra High on fast mode plus Computer Use beats Opus 5.5 — soumitrashukla9 · 2026-09-26
- Vercel launches stealth model Pixel Canary, ties GPT-6 Astra on Next.js evals — cramforce · 2026-09-26
- GPT-6-Astra-Max Tops BALROG Game Benchmark at 68.3%, NetHack Depth Hits 13.2 — burny_tech · 2026-09-26