Kimi K3 Leads in Autonomous Legal Work Benchmark
rohanpaul_ai · x · 2026-07-19
In a rigorous benchmark test for autonomous legal work, Kimi K3 took a massive lead with a 26.7% success rate, far outpacing the runner-up Claude Fable 5 (14.2%).
The test comprises 120 real-world tasks across 24 legal domains (such as drafting memos and discovery summaries), requiring models to autonomously read case files and generate complete legal documents. Grading is exceptionally strict—missing any single detail results in a failed task, explaining why even the leader's success rate sits at only 26.7%.
Overall, completing only about 27 out of 100 tasks indicates that AI still has a long way to go in fully unsupervised legal work.
Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→
More from Models
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22