Grok 4.6 tops agentic benchmark, halves cost vs Claude rival
elonmusk · x · 2026-08-18
Elon Musk shared data showing Grok 4.6 ties for first place on the Artificial Analysis Agentic Index with a score of 59, alongside Claude Opus 5 Max. The index evaluates tool use, planning, and autonomy. Grok 4.6 completes tasks in 53 turns using 0.5bn input tokens on average, significantly more efficient than Claude's 103 turns and 2.0bn tokens. At $0.84 per task, Grok achieves optimal intelligence-to-cost efficiency.
More from coding & agent
- Researcher: Agents that can work in a sandbox shouldn't have outbound access at all — moniquejmorrow · 2026-10-02
- llama.cpp adds decision models: typed questions, per-option probabilities in one forward pass — ngxson · 2026-10-02
- Agents aren't GPU-bound: tool execution, memory bandwidth and sandbox overhead are the real bottleneck — ai · 2026-10-02
- Ruff creator: 'no one writes code anymore' isn't the point — does anyone read it? — charliermarsh · 2026-10-02
- 8 Practical Upgrades to Turn Fragile Chat Loops Into Production-Grade Agents — alexcovo_eth · 2026-10-02
- Event-Driven Agents Are Right for Real Work, But AWS Stack Makes Humans the Bottleneck — TansuYegen · 2026-10-02