14 个用例复现完整榜单:Terminal Bench Mini 加速本地模型评测

asankhs · reddit · 2026-09-22

Reddit 用户 asankhs 发布了 Terminal Bench Mini,一个仅含 14 个实例的 terminal-bench 子集,参考 deepswe-mini 的思路构建。它能在本地 LLM 上以远快于完整 agentic 评测的速度运行,同时生成与 terminal-bench 完整排行榜高度一致的代表性排名,解决本地模型跑标准 agent 评测太慢的痛点。数据集已在 Hugging Face 上开源(LocalLLaMA/terminal-bench-mini)。

原文链接 →

「研究」频道最新

更多「研究」频道 AI 资讯 →