Grok 4.5 Tops Real-World Task Benchmark
XFreeze · x · 2026-07-11
Grok 4.5 secured #1 on AutomationBench-AA, which measures real-world AI automation capabilities.
- Scored 51%, beating Claude Fable 5 (49%) and Claude Opus 4.8 (48%)
- Claims to cost roughly 1/4 of competitors per task
- Uses only about 8K output tokens per task, which the author notes is highly token-efficient for a frontier model
The post emphasizes that Grok 4.5 isn't just leading in scores, but also stands out in cost and token efficiency for real-world tasks.
More from coding & agent
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- Claude Code skill uses 10 Markdown rules to make outputs ADHD-friendly — alex_verem · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22
- A better path to agent autonomy is running waves, finding friction, and iterating — JnBrymn · 2026-07-22
- AI agent designers map the visual and tonal cues behind companionship products — Unlikely-Platform-47 · 2026-07-22
- Coding agents are heading toward an AI-writes, AI-reviews, human-approves workflow — aftahi_ai · 2026-07-22