Artificial Analysis grew from 4 exam-style evals to 10, adding long-horizon agent tasks in two years
davidyin44 · x · 2026-09-26
A user observed that Artificial Analysis expanded from 4 exam-style evals to 10 benchmarks in two years, now including long-horizon agent tasks — making old leaderboards hard to compare with the current one and illustrating how model evaluation standards are rapidly evolving.
More from Models
- Dev swaps DeepSeek Flash for Opus 5.5 to fix endless bug loops — oran_ge · 2026-09-26
- Game built by Opus 5.5 syncs real NYC weather—NPCs carry umbrellas in the rain — mattshumer_ · 2026-09-26
- Dev says Opus 5.5 ignored explicit instructions, bundling the wrong app: 'they lobotomized Opus' — haydendevs · 2026-09-26
- Open-weights AI is now mostly Chinese labs, weekly token share data shows — FinanceYF5 · 2026-09-26
- Building Speech AI Book Hits Kindle: Spectrograms to Real-Time Voice Agents, Code in Every Chapter — prdeepakbabu · 2026-09-26
- User flags Claude weekly usage: 5% gone before first session even ends — ColleenMBrady · 2026-09-26