Kimi K3 Leads in Autonomous Legal Work Benchmark

rohanpaul_ai · x · 2026-07-19

In a rigorous benchmark test for autonomous legal work, Kimi K3 took a massive lead with a 26.7% success rate, far outpacing the runner-up Claude Fable 5 (14.2%).

The test comprises 120 real-world tasks across 24 legal domains (such as drafting memos and discovery summaries), requiring models to autonomously read case files and generate complete legal documents. Grading is exceptionally strict—missing any single detail results in a failed task, explaining why even the leader's success rate sits at only 26.7%.

Overall, completing only about 27 out of 100 tasks indicates that AI still has a long way to go in fully unsupervised legal work.

Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→

Original post →

More from Models

Models channel →