Kimi K3 Leads Legal Benchmark
petrusenko_max · x · 2026-07-19
Kimi K3 scored **26.7%** on a challenging **autonomous legal work benchmark**, nearly doubling the score of its recent competitor Claude Fable 5 (**14.2%**). The test covers 24 legal areas with 120 private tasks, including memos and deposition summaries. Models are required to fully process case files and output final documents under strict grading: missing any single rubric detail results in a failure. The author points out that while 27% is a leading score, there is still a long way to go before AI can reliably and independently handle legal work.
Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→
More from Models
- Qwen 3.8 Max, Kimi K3, OpenAI’s long-horizon bug, and a crowded infra day — Latent Space · 2026-07-21
- Kimi K3 Max beats Qwen 3.8 Max preview on sketch-to-3D and Strandbeest tasks — OfirPress · 2026-07-21
- Kimi post pairs a compute anecdote with a subjective top-5 model ranking — vista8 · 2026-07-21
- DavidAU collaboration reportedly improves Qwen 3.6 27B on long-context agent tasks — My_Unbiased_Opinion · 2026-07-21
- Smaller language models stay terse while larger ones explain the physics — NickPassig · 2026-07-21
- Sources and Lean proof links for the IMO 2026 model benchmark — deedydas · 2026-07-21