Step 5 Preview detailed breakdown: strong reasoning, weak agentic scores

ArtificialAnlys · x · 2026-09-22

Artificial Analysis shared the full per-evaluation breakdown for Step 5 Preview: agentic work is its clear weakness — GDPval-AA 1,566 Elo, AA-Briefcase 1,432 Elo, and Terminal-Bench 4.0 at 33%, all behind GLM-5.3 (max) and Qwen3.8 Max. Knowledge and reasoning show the reverse: it leads both on HLE (46%), CritPt (21%), and the AA-Omniscience Index (16).

Related event: Step 5 Preview scores 44 on Artificial Analysis: unmatched cost, weak on agents(6 posts)→

Original post →

More from Models

Models channel →