Kimi K3 Reportedly Scores Twice as High as Claude 3.5 on Autonomous Legal Tasks
togethercompute · x · 2026-08-07
According to Together Compute, Kimi K3 scores nearly twice as high as Claude Fable 5 (likely a typo or alias for Claude 3.5) on Harvey LAB-AA's hard autonomous legal tasks.
Related event: Kimi K3 Nearly Doubles Claude's Score in Complex Legal Tasks(2 posts)→
More from Models
- Open Video Models Catch the Frontier: MiniMax H3, FLUX 3, and Seedance 2.5 — altryne · 2026-08-07
- Claude Fabricates Numbers Instead of Admitting Uncertainty, Frustrating Users — k1_r1 · 2026-08-07
- mLateOn Hits SOTA on MTEB Korean Retrieval Without Any Korean Training Data — IgorCarron · 2026-08-07
- Just 0.3% Behind? Netizens Mock AI Benchmark Marketing Spin — teortaxesTex · 2026-08-07
- Cognition & OpenRouter on Model Routing: Why Naive Task Routing Fails for Agents — AI Engineer · 2026-08-07
- MiniMax-H3 Base vs. Turbo-LoRA Version Compared — _akhaliq · 2026-08-07