Kimi K3 Leads Legal Benchmark

petrusenko_max · x · 2026-07-19

Kimi K3 scored 26.7% on a challenging autonomous legal work benchmark, nearly doubling the score of its recent competitor Claude Fable 5 (14.2%).

The test covers 24 legal areas with 120 private tasks, including memos and deposition summaries. Models are required to fully process case files and output final documents under strict grading: missing any single rubric detail results in a failure. The author points out that while 27% is a leading score, there is still a long way to go before AI can reliably and independently handle legal work.

Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→

Original post →

More from Models

Models channel →