105 个真实 bug 实测:Mistral Large 4 惨败 Qwen 与 DeepSeek

PawelHuryn · x · 2026-10-07

Pawel Huryn 用自建的 Bug Hunt Bench(两个真实仓库、105 个植入 bug、盲评、每仓库一次提示)测试 Mistral Large 4 与其他前沿编程模型,结果 Mistral 排名垫底:

结论:Mistral Large 4 不是最强、不是最便宜也不是最快。榜单所有运行日志、成本、轮次数据都在配套网站上公开,可按成本/时间/轮次排序对比,细节与注意事项在 GitHub 的 run-notes 中。

原文链接 →

「模型」频道最新

更多「模型」频道 AI 资讯 →