Grok 4.5 Tops Real-World Task Benchmark
XFreeze · x · 2026-07-11
Grok 4.5 secured #1 on AutomationBench-AA, which measures real-world AI automation capabilities.
- Scored 51%, beating Claude Fable 5 (49%) and Claude Opus 4.8 (48%)
- Claims to cost roughly 1/4 of competitors per task
- Uses only about 8K output tokens per task, which the author notes is highly token-efficient for a frontier model
The post emphasizes that Grok 4.5 isn't just leading in scores, but also stands out in cost and token efficiency for real-world tasks.
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11