Musk Amplifies Claim: Grok 4.7 Scores 19.6% on Legal Agent Benchmark, Nearly 3x Fable 5.1
elonmusk · x · 2026-09-22
Elon Musk retweeted a claim that Grok 4.7 tops the Legal Agent Benchmark, scoring 19.6% on realistic, long-term legal tasks — beating every other model tested and nearly 3x the score of Fable 5.1.
Note this comes from a third-party fan account rather than an official source, and the low absolute score shows how hard long-horizon legal agent tasks remain; independent verification is still pending.
More from Models
- Rumored OpenAI reasoning mode "Aeon" would persist until the task is done — imjustnewatai · 2026-09-22
- 105 Real Bugs Tested: Grok 4.7 Scores 28.7, No Better Than Grok 4.6; GPT-6 Astra Leads at 45 — PawelHuryn · 2026-09-22
- OpenAI's rumored 'Aeon' is a persistent reasoning mode that runs until the task is done — imjustnewatai · 2026-09-22
- Grok 4.7 vs Claude Fable 5.1 vs GPT-6 Astra vs DeepSeek V4.1 Flash: 40+ benchmark showdown — HealthySkeptic2000 · 2026-09-22
- Grok 4.7 xHigh hits 58% on AA-Briefcase, just 1 point behind Claude Fable 5.1 Max — XFreeze · 2026-09-22
- Gemini beats GPT-6 Astra at robot capture the flag, winning 70% of matches — chris_j_paxton · 2026-09-22