Stronger Models Take Longer on Agent Tasks
ArtificialAnlys · x · 2026-07-08
Data from Artificial Analysis indicates that stronger models tend to spend more time per task: Claude Fable 5 takes an average of 16.9 minutes/task, Claude Opus 4.8 18.5 minutes, and Claude Sonnet 5 22.8 minutes. GLM-5.2 is an exception, achieving a full pass rate comparable to Opus 4.8 in only 5.0 minutes/task. DeepSeek V4 Flash is the fastest model to reach a non-zero full pass rate, averaging 4.4 minutes.
Related event: Artificial Analysis Releases Harvey Legal Agent Benchmark Results(8 posts)→
More from Models
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22
- Gemini 3.5 Flash-Lite beats 3.1 Flash-Lite on long-context retrieval in MRCRv2 — Dillonu · 2026-07-22
- Poolside launches Laguna S 2.1, a 118B MoE with a 1M-token context window — Lowkey_LokiSN · 2026-07-22