Kimi K3 Max Agent Benchmark: Delivers 2.8x More Solved Tasks Per Dollar

togethercompute · x · 2026-08-03

Together Compute released an analysis of the Kimi K3 Max model on the DeepSWE benchmark. Data shows that Kimi K3 Max achieves Pass@1 performance close to Fable 5 xhigh, but at approximately one-third the cost per rollout.

This results in 2.8x more solved tasks per dollar. For teams running models at scale for agentic workflows, the cost per successful task is often a more practical metric for comparison.

Original post →

More from coding & agent

coding & agent channel →