Step 5 Preview beats GLM-5.3 on factual recall but hallucinates on 43% of attempts
ArtificialAnlys · x · 2026-09-22
Artificial Analysis notes Step 5 Preview hits 42% accuracy on AA-Omniscience with only 600B total parameters, ahead of GLM-5.3 (max, 34% at 753B) and slightly above the parameter-count scaling trend, though behind Kimi K3 (max, 48% at 2.8T). The limiter is behavior rather than knowledge: it attempts 68% of questions and hallucinates on 43% of those attempts instead of declining to answer.
More from Models
- xAI's Grok hype cycle repeats: 4.7 delayed, no frontier model beats ChatGPT or Claude — flowersslop · 2026-09-22
- xAI fixes SDK bug dropping reasoning content, significantly boosting Grok 4.7 — ns123abc · 2026-09-22
- TTS leaderboard: xAI hits 87.6% pronunciation accuracy, Kokoro 82M fastest at 242 chars/s — ArtificialAnlys · 2026-09-22
- xAI's Post-Launch SDK Update Pushes It to #10 on the Vals Index — teortaxesTex · 2026-09-22
- Grok 4.7 looks pricier than before, now costlier than Astra on Artificial Analysis — steipete · 2026-09-22
- Which LLM is most encyclopedic on 8GB VRAM + 64GB RAM? — Mangleus · 2026-09-22