DeepSeek V4 Flash beats Claude via self-verification on Terminal-Bench

nptacek · x · 2026-08-21

Sampling 5 solutions with DeepSeek V4 Flash and ranking them using the same model (LLM-as-a-Verifier) boosts accuracy on Terminal-Bench 2.1 from 79% to 88%, outperforming closed frontier models at 11x lower cost.

Related event: Stanford's Open Framework Helps DeepSeek V4 Flash Outperform Claude(2 posts)→

Original post →

More from coding & agent

coding & agent channel →