StepFun's Step 5 Preview scores 44 on AA Intelligence Index at ~2.8x lower cost than Kimi K3
ArtificialAnlys · x · 2026-09-22
Artificial Analysis published a full evaluation of StepFun's new flagship Step 5 Preview, scoring 44 on the Intelligence Index — matching Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45).
Highlights
- Cost edge: $0.72 per Intelligence Index task, 2.8x cheaper than Kimi K3 (max) at the same score ($2.00), driven by $1/$2.70 per 1M input/output token pricing. Only MiMo-V2.6-Pro scores higher (46) at a lower cost ($0.13).
- Big reasoning jump: 46% on Humanity's Last Exam (+25 pts vs Step 3.7 Flash) and 21% on CritPt (+19), roughly in line with Kimi K3.
- Knowledge vs. hallucination: 42% accuracy on AA-Omniscience at 600B total params, beating GLM-5.3 (753B), but it hallucinates on 43% of attempts instead of declining, capping its index at 16.
- Agentic weakness: lags peers on GDPval-AA (1,566 Elo), Terminal-Bench 4.0 (33%), and AA-Briefcase (1,432 Elo).
Specs: 600B total / 27B active MoE, 1M context, text/image/video input; closed-weights API now, open weights planned for October 15.
More from Infra
- Cerebras CEO on Jensen Huang: a decade trading as 'nobody' before Nvidia made it — rohanpaul_ai · 2026-09-22
- How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization — TencentHunyuan · 2026-09-22
- Formal CRC guarantee for a semantic cache breaks when the cache's own writes shape future calibration data — Reasonable_Royal_621 · 2026-09-22
- 456GB DeepSeek v4.1 runs locally at 40 tok/s with Threadripper + dual RTX 6000 hybrid setup — HankYeomans · 2026-09-22
- Agent Substrate roadmap: sub-second suspend/resume runtime for dense agent deployments — rakyll · 2026-09-22
- SemiAnalysis tears down A20 on TSMC N2, DRAM revealed in iPhone 18 Pro Max package — dylan522p · 2026-09-22