Kimi K3 Leads Legal Benchmark

petrusenko_max · x · 2026-07-19

Kimi K3 scored **26.7%** on a challenging **autonomous legal work benchmark**, nearly doubling the score of its recent competitor Claude Fable 5 (**14.2%**). The test covers 24 legal areas with 120 private tasks, including memos and deposition summaries. Models are required to fully process case files and output final documents under strict grading: missing any single rubric detail results in a failure. The author points out that while 27% is a leading score, there is still a long way to go before AI can reliably and independently handle legal work.

Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→

Original post →

More from Models

Models channel →