Terminal-bench 榜单:Kimi 夺冠,DeepSeek 第二

SergioPaniego · x · 2026-08-24

Terminal-bench 被认为是去年最重要的基准测试之一,包含 89 个高质量任务(如修复 OCaml GC bug、恢复 SQLite 数据库),用于指导本地 Agent 的高效发展。

最新排行榜:

原文链接 →

「模型」频道最新

更多「模型」频道 AI 资讯 →