Step 5 Preview detailed breakdown: strong reasoning, weak agentic scores
ArtificialAnlys · x · 2026-09-22
Artificial Analysis shared the full per-evaluation breakdown for Step 5 Preview: agentic work is its clear weakness — GDPval-AA 1,566 Elo, AA-Briefcase 1,432 Elo, and Terminal-Bench 4.0 at 33%, all behind GLM-5.3 (max) and Qwen3.8 Max. Knowledge and reasoning show the reverse: it leads both on HLE (46%), CritPt (21%), and the AA-Omniscience Index (16).
More from Models
- xAI's Grok hype cycle repeats: 4.7 delayed, no frontier model beats ChatGPT or Claude — flowersslop · 2026-09-22
- xAI fixes SDK bug dropping reasoning content, significantly boosting Grok 4.7 — ns123abc · 2026-09-22
- TTS leaderboard: xAI hits 87.6% pronunciation accuracy, Kokoro 82M fastest at 242 chars/s — ArtificialAnlys · 2026-09-22
- xAI's Post-Launch SDK Update Pushes It to #10 on the Vals Index — teortaxesTex · 2026-09-22
- Grok 4.7 looks pricier than before, now costlier than Astra on Artificial Analysis — steipete · 2026-09-22
- Which LLM is most encyclopedic on 8GB VRAM + 64GB RAM? — Mangleus · 2026-09-22