DeepSeek V4 tested: ARC-AGI reasoning costs drop despite higher performance

DeArgonaut · reddit · 2026-08-09

A developer tested DeepSeek V4 (0731) on the ARC-AGI-1 and ARC-AGI-2 benchmarks. The results show that this version breaks the industry trend where performance gains come with increased reasoning costs, achieving higher scores while actually decreasing the cost per task.

Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance, Inference Cost Drops as Performance Rises(5 posts)→

Original post →

More from Models

Models channel →