Stanford's Self-Verification Boosts DeepSeek Past Claude

机器之心 · wechat · 2026-08-27

Result: Stanford researchers introduced LLM-as-a-Verifier, a framework where open-source models (e.g., DeepSeek V4 Flash) generate and verify their own candidates. This boosted success rates on Terminal-Bench V2 from 79% to 88%, at roughly 1/11th the cost of Claude Fable 5.

Methodology:

Applications:

Efficiency: Reduces selection complexity from O(N²) to O(Nk) using a pivot-based algorithm.

Original post →

More from coding & agent

coding & agent channel →