Kimi K3 Leads in Autonomous Legal Work Benchmark
rohanpaul_ai · x · 2026-07-19
In a rigorous benchmark test for autonomous legal work, Kimi K3 took a massive lead with a 26.7% success rate, far outpacing the runner-up Claude Fable 5 (14.2%).
The test comprises 120 real-world tasks across 24 legal domains (such as drafting memos and discovery summaries), requiring models to autonomously read case files and generate complete legal documents. Grading is exceptionally strict—missing any single detail results in a failed task, explaining why even the leader's success rate sits at only 26.7%.
Overall, completing only about 27 out of 100 tasks indicates that AI still has a long way to go in fully unsupervised legal work.
Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- giffmana: the env being used in training is part of the point — giffmana · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11