Kimi K3 Leads Legal Benchmark
petrusenko_max · x · 2026-07-19
Kimi K3 scored 26.7% on a challenging autonomous legal work benchmark, nearly doubling the score of its recent competitor Claude Fable 5 (14.2%).
The test covers 24 legal areas with 120 private tasks, including memos and deposition summaries. Models are required to fully process case files and output final documents under strict grading: missing any single rubric detail results in a failure. The author points out that while 27% is a leading score, there is still a long way to go before AI can reliably and independently handle legal work.
Related event: Kimi K3 Leads Harvey Legal Benchmark(4 posts)→
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11