Gemini 3.8 Flash 跑 DeepSWE 得 73.7%,领先 Fable 5.1 六个百分点
cedric_chee · x · 2026-09-03
Gemini 3.8 Flash 在 DeepSWE 基准上取得 73.7% 的成绩,相比之下 Fable 5.1 为 67.4%,体现出谷歌新 Flash 模型在自主软件工程任务上的优势。
所属事件:Gemini 3.8 Flash 影子发布:跑分曝光、定价持平、3.5 Pro 被弃(42 条相关)→
「模型」频道最新
- Google 六周内第三款 Flash:Gemini 3.8 Flash 强化 Agent 与编码 — brianryhuang · 2026-09-03
- 俄罗斯团队创 Mostik:让模型用权重「心电感应」,混合模型登顶 ARC-AGI 3 — nordicinst · 2026-09-03
- Anthropic 上线 Claude 文件检测工具,C2PA 凭证一键验证出处 — btibor91 · 2026-09-03
- Marin 535B A23B 训练公开博客与 wandb 指标入口 — Sentdex · 2026-09-03
- 字节开源 LoopLM:循环潜空间推理让 1.4B 模型比肩 12B 竞品,Bengio 参与 — peterjliu · 2026-09-03
- Marin 535B A23B 前沿级训练全程实时公开可追踪 — Sentdex · 2026-09-03