Grok 4.6 ranks #2 in legal agent eval at 13× lower cost than GPT-5.6

XFreeze · x · 2026-08-18

Grok 4.6 achieved a #2 ranking in agentic U.S. legal research benchmarks with 62.12% accuracy. It slightly trails Claude Opus 5 (65.15%) but beats GPT-5.6 Sol (60.61%). The standout is cost efficiency: at $1.52 per test, it is roughly 13× cheaper than GPT-5.6 ($19.69) and over 4× cheaper than Opus ($6.58), delivering frontier-level performance at a fraction of the cost.

Original post →

More from Models

Models channel →