SlopCodeBench 横评:GPT-6 Astra 21% 领跑,GLM 5.3 High 达 19%

teropa · x · 2026-09-06

dexhorthy 更新自建的 SlopCodeBench 编码基准,首次在多模型上按不同思考强度(effort 设置)对比:GPT-6 Astra Xhigh 得分 21%(仍在跑)、GPT-5.6 Sol XHigh 18%(进行中)、Fable 5.1 Medium 11%、GPT-5.6 Sol medium 11%、GLM 5.3 Medium 7%、GLM 5.3 High 19%(进行中)。作者指出思考强度不能在不同厂商之间直接换算;因 Fable 额度 4 天后重置,除非有人赞助额度,Fable 5.1 xhigh 只能等下周再测。

原文链接 →

「模型」频道最新

更多「模型」频道 AI 资讯 →