Scale Releases MCP Tool-Calling Eval, Kimi K3 Tops Open-Source Models
LiTianleli · x · 2026-07-31
Scale Labs updated its MCP Atlas leaderboard, a benchmark designed to evaluate LLMs on realistic, multi-step tool use via the Model Context Protocol (involving 1,000 tasks and 36 MCP servers).
In the latest results, Kimi K3 ranks first among open-source models and outperforms Gemini 3 flash lite and GPT 5.6 luna in long-horizon tool calling.
More from Models
- Luna Offers Cheaper Inputs, but DeepSeek Wins Cache Economics in Long Agentic Sessions — teortaxesTex · 2026-07-31
- User Criticizes Hive AI: Fake Image Missed, Fake Sound Detected in Silent Clip — henkvaness · 2026-07-31
- DeepSWE Coding Agent Leaderboard: Claude Opus and GPT-5.6 Top the Charts — tristanbob · 2026-07-31
- Hive AI Detection Fails: Fake Crows Missed, 36% Music Detected in Silent Clip — henkvaness · 2026-07-31
- Google Responds to AI Misinformation Concerns: Gemini Images Embed SynthID Watermarks — henkvaness · 2026-07-31
- Users report OpenAI's o1-pro model got slower but smarter — teortaxesTex · 2026-07-31