Step 5 Preview beats GLM-5.3 on factual recall but hallucinates on 43% of attempts

ArtificialAnlys · x · 2026-09-22

Artificial Analysis notes Step 5 Preview hits 42% accuracy on AA-Omniscience with only 600B total parameters, ahead of GLM-5.3 (max, 34% at 753B) and slightly above the parameter-count scaling trend, though behind Kimi K3 (max, 48% at 2.8T). The limiter is behavior rather than knowledge: it attempts 68% of questions and hallucinates on 43% of those attempts instead of declining to answer.

Related event: Step 5 Preview scores 44 on Artificial Analysis: unmatched cost, weak on agents(6 posts)→

Original post →

More from Models

Models channel →