Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark
9 月 2 日至 3 日,Google 的 Logan Kilpatrick 分享了新模型 Gemini 3.8 Flash 在编程/Agent 基准 DeepSWE 1.1 上的成绩:73.7%,随后多位用户在社交平台转述并补充了对比数据,引发关注。
已确认
- Logan Kilpatrick(Google)公布了 Gemini 3.8 Flash 在 DeepSWE 1.1 上的成绩 73.7%,并附评测链接(m1、m4)。
- @cedricchee 补充对比数据:Gemini 3.8 Flash 得 73.7%,Fable 5.1 为 67.4%,领先约六个百分点(m3)。
- @cedricchee 还列出 Gemini 3.8 Flash(high)pass@1 达 73.7%,与旗舰模型几乎持平:Opus 5(max)为 74.0%(pass@1),并提到 GPT-5.6 等模型(帖子原文被截断,完整对比数据需查看原帖)(m5)。
尚未确认
- @vedantmisra 提醒,该消息目前主要是第三方转述与截图,Google 官方尚未正式发布确认,评测细节需点开原链接核实(m2)。
为什么重要
- 若成绩属实,Flash 级模型在自主软件工程任务上已逼近甚至追平旗舰模型,性价比意义明显。
- @doodlestein 表示期待:3.7 Flash 已经不错,3.8 Flash 进一步提升更值得观望(m4)。
2026-09-02 ~ 2026-09-03 · 5 related posts
Primary sources
- Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark — OfficialLoganK ·
- Gemini 3.8 Flash scores 73.7% on DeepSWE, beating Fable 5.1 by 6 points — cedric_chee ·
- Gemini 3.8 Flash hits 73.7% on DeepSWE, nearly matching Opus 5 and GPT-5.6 Sol — cedric_chee ·
- [source] Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark — OfficialLoganK · 2026-09-02
- [source] Gemini 3.8 Flash scores 73.7% on DeepSWE, beating Fable 5.1 by 6 points — cedric_chee · 2026-09-03
- [source] Gemini 3.8 Flash hits 73.7% on DeepSWE, nearly matching Opus 5 and GPT-5.6 Sol — cedric_chee · 2026-09-03
- Gemini 3.8 Flash posted scoring 73.7% on DeepSWE 1.1 — doodlestein · 2026-09-03
- Unverified claim: Gemini 3.8 Flash tops DeepSWE v1.1 coding benchmark — vedantmisra · 2026-09-03