斯坦福开源框架助 DeepSeek V4 Flash 反超 Claude
斯坦福团队为 DeepSeek V4 Flash 推出开源验证框架,通过先生成 5 套方案再以同一模型作为 LLM-as-a-Judge 排序筛选最优解的自验证机制,将其在 Terminal-Bench 2.1 上的准确率从 79% 进一步提升,成功反超 Claude,而成本仅为后者的约 1/11,展示了推理时采样与自评策略的性价比。
2026-08-20 ~ 2026-08-21 · 2 条相关
- 斯坦福框架助 DeepSeek 超越 Claude,成本仅 1/11 — FuSheng_0306 · 2026-08-20
- DeepSeek V4 Flash 靠自验证击败 Claude — nptacek · 2026-08-21