Mistral Large 4 榜单落后 GLM-5.3,Yuchen 调侃「评测危机」

Yuchenj_UW · x · 2026-10-06

AI 研究者 Yuchen Jin 指出,Mistral 刚发布的 Mistral Large 4(绰号 Le Chonk)在 Artificial Analysis Intelligence Index 上明显落后于 GLM-5.3,甚至不及 GLM-5.3 Flash。他借此调侃当前评测基准乱象,称「我们正处于 Eval 危机中」——不同榜单结果与各家宣传口径差异巨大,模型真实能力越来越难判断。

原文链接 →

「Fun」频道最新

更多「Fun」频道 AI 资讯 →