Step 5 Preview scores 44 on Artificial Analysis: unmatched cost, weak on agents

On September 22, Artificial Analysis published its full evaluation of StepFun's new flagship Step 5 Preview: the model scored 44 overall, on par with Kimi K3 (max), slightly behind GLM-5.3 (max, 45) and Qwen3.8 Max, ranking 26th among 202 models on the leaderboard. Overall, it's a flagship whose biggest selling point is extremely low pricing, with leading reasoning ability but comparatively weak agent capabilities.

Confirmed

Why it matters

Step 5 Preview charts a differentiated path: rather than winning on absolute scores, it targets the value-for-money segment with 44-point intelligence at roughly 1/2.8 the per-task cost. But its 43% hallucination rate and agent weaknesses suggest caution is warranted for automation scenarios requiring high reliability.

2026-09-22 ~ 2026-09-22 · 6 related posts

Full story(2 episodes)→

Primary sources