Grok 4.5 Tops Multiple Professional AI Benchmarks

According to the latest AI knowledge work benchmarks released by Artificial Analysis, Grok 4.5 has demonstrated strong professional capabilities, becoming the best-performing non-Anthropic model currently available. This achievement has drawn attention within the AI community, proving its competitiveness in complex tasks.

Key Details

Grok 4.5 has achieved leading scores in multiple professional work benchmarks. According to @kevinnbass, the model's scores significantly outperform other frontier models such as Opus 4.8 and GPT 5.5. Detailed data from @XFreeze points out that Grok 4.5 took first place in several细分 leaderboards, including AutomationBench-AA, Terminal-Bench v2, Harvey Legal Agent Benchmark, SWE Marathon, and SWE-Atlas.

Areas of Excellence

Multiple authors specifically highlighted Grok 4.5's outstanding performance in legal tasks. Whether in comprehensive evaluations or the specific Harvey Legal Agent Benchmark, the model has shown a prominent capacity for handling professional legal work.

2026-07-10 ~ 2026-07-10 · 5 related posts

2 near-duplicate retellings: kevinnbass · kevinnbass