FULL STORY

Step 5 Preview: Launch and First Benchmarks

Stepfun launched Step 5 Preview on Sept 21 with a cost-performance positioning. Artificial Analysis then benchmarked it at 44 points, noting strong cost efficiency but agent weaknesses.

2026-09-21 ~ 2026-09-22 · 2 episodes · 9 posts

Episode 1 · StepFun launches Step 5 Preview, touting cost-performance frontier (2026-09-21, 3 posts)

StepFun released its flagship Step 5 Preview, positioned on the cost-performance Pareto frontier; early hands-on testing found it beats GLM 5.3 on long tasks and rivals GLM and Kimi flagships on value.

Episode 2 · Step 5 Preview scores 44 on Artificial Analysis: unmatched cost, weak on agents (2026-09-22, 6 posts)

On September 22, Artificial Analysis published its full evaluation of StepFun's new flagship Step 5 Preview: the model scored 44 overall, on par with Kimi K3 (max), slightly behind GLM-5.3 (max, 45) and Qwen3.8 Max, ranking 26th among 202 models on the leaderboard. Overall, it's a flagship whose biggest selling point is extremely low pricing, with leading reasoning ability but comparatively weak agent capabilities.

Confirmed

  • Specs and overall score: 600B total parameters / 27B activated, Intelligence Index of 44, ranked 26th; output speed of 92.8 tokens/sec, faster than the 71 average; priced at $1/$2.70 per million tokens, both below the median.
  • Reasoning and agent performance: leading reasoning, but across-the-board weakness on agent tasks is a clear shortcoming—GDPval-AA 1,566 Elo, AA-Briefcase 1,4xx (the post didn't give the full number).
  • Knowledge and hallucination: AA-Omniscience accuracy of 42%, beating GLM-5.3 (max, 753B, 34%) and slightly above the factual-recall trend line for its parameter count; however, 43% of its answers contained hallucinations.
  • Cost advantage: running the full Intelligence Index costs about $0.72 per task, versus roughly $2.00 for the same-scoring Kimi K3 (max) and even more for the 1-point-higher GLM-5.3 (max); Artificial Analysis notes it breaks the cost line through pricing rather than token efficiency, at roughly 1/2.8 the cost of peer models.

Why it matters

Step 5 Preview charts a differentiated path: rather than winning on absolute scores, it targets the value-for-money segment with 44-point intelligence at roughly 1/2.8 the per-task cost. But its 43% hallucination rate and agent weaknesses suggest caution is warranted for automation scenarios requiring high reliability.