Musk Amplifies Claim: Grok 4.7 Scores 19.6% on Legal Agent Benchmark, Nearly 3x Fable 5.1

elonmusk · x · 2026-09-22

Elon Musk retweeted a claim that Grok 4.7 tops the Legal Agent Benchmark, scoring 19.6% on realistic, long-term legal tasks — beating every other model tested and nearly 3x the score of Fable 5.1.

Note this comes from a third-party fan account rather than an official source, and the low absolute score shows how hard long-horizon legal agent tasks remain; independent verification is still pending.

Original post →

More from Models

Models channel →