Grok 4.6 tops agentic benchmark, halves cost vs Claude rival
elonmusk · x · 2026-08-18
Elon Musk shared data showing Grok 4.6 ties for first place on the Artificial Analysis Agentic Index with a score of 59, alongside Claude Opus 5 Max. The index evaluates tool use, planning, and autonomy. Grok 4.6 completes tasks in 53 turns using 0.5bn input tokens on average, significantly more efficient than Claude's 103 turns and 2.0bn tokens. At $0.84 per task, Grok achieves optimal intelligence-to-cost efficiency.
Related event: Grok 4.6 Tops Agentic Index at a Fraction of Rivals' Cost(2 posts)→
More from coding & agent
- Gitlawb outperforms Cursor in benchmark, 36/s vs 24/s commit rate — JoshuaJBouw · 2026-08-18
- Tidebroker: Secure Credential Brokerage for AI Agents — steipete · 2026-08-18
- Robotics demo: 3D metal dog built entirely with Grok-generated code — techartist_ · 2026-08-18
- Perplexity Computer adds manual intervention for agent tools — AravSrinivas · 2026-08-18
- Forcing Codex to optimize for 10 hours achieved 200× speedup — josh_wills · 2026-08-18
- Concept: LLM memory management via a structured file system — Orectoth · 2026-08-18