Kimi K3 scores 60% higher than Fable 5.1 on Harvey's hard autonomous legal tasks
togethercompute · x · 2026-09-10
Together Compute reports that Kimi K3 scores 60% higher than Fable 5.1 on Harvey lab-aa's hard autonomous legal task benchmark — an eval focused on real-world legal workflows requiring autonomous task completion rather than single-turn Q&A. If it holds up, it marks a significant vertical-domain capability jump for the new Kimi release.
More from Models
- Pixel analysis of OpenAI's Navier-Stokes chart suggests ~91 solved open problems kept unreleased — teortaxesTex · 2026-09-10
- Developer who spent €20k+ on Claude says models are degrading, urges skipping yearly plans — maier_ak · 2026-09-10
- User who spent €20k+ on Claude licenses says models keep getting weaker — maier_ak · 2026-09-10
- Users report AI assistant Astra omitting spaces, most often around numbers — gandamu_ml · 2026-09-10
- DeepSeek V4.1 flash reportedly has hidden vision support; official release still missing — teortaxesTex · 2026-09-10
- GPT-6 Astra vs GPT-5.6 Sol: Code Review Benchmark on 50 Real PRs — entelligenceai17 · 2026-09-10