Kimi K3 scores 60% higher than Fable 5.1 on Harvey's hard autonomous legal tasks

togethercompute · x · 2026-09-10

Together Compute reports that Kimi K3 scores 60% higher than Fable 5.1 on Harvey lab-aa's hard autonomous legal task benchmark — an eval focused on real-world legal workflows requiring autonomous task completion rather than single-turn Q&A. If it holds up, it marks a significant vertical-domain capability jump for the new Kimi release.

Original post →

More from Models

Models channel →