Scale Releases MCP Tool-Calling Eval, Kimi K3 Tops Open-Source Models

LiTianleli · x · 2026-07-31

Scale Labs updated its MCP Atlas leaderboard, a benchmark designed to evaluate LLMs on realistic, multi-step tool use via the Model Context Protocol (involving 1,000 tasks and 36 MCP servers).

In the latest results, Kimi K3 ranks first among open-source models and outperforms Gemini 3 flash lite and GPT 5.6 luna in long-horizon tool calling.

Original post →

More from Models

Models channel →