DeepSeek V4 Flash beats Claude on Terminal-Bench with self-verification at 1/11 cost

ccerrato147 · x · 2026-08-18

By enabling self-verification during inference, DeepSeek V4 Flash outperforms Claude Fable on Terminal-Bench 2.1 while being 11x cheaper. The method involves sampling 5 candidates and ranking them using LLM-as-a-Verifier, boosting accuracy from 79% to 88%.

Related event: DeepSeek V4 Flash Beats Claude Fable 5 on Terminal-Bench via Self-Verification at 1/11 the Cost(5 posts)→

Original post →

More from Research

Research channel →