DeepSeek V4 Flash Scores Impressively on ARC-AGI Semi-Private Tasks

petrusenko_max · x · 2026-08-09

According to a user's test, the DeepSeek V4 Flash 0731 model achieved 89.0% on ARC-AGI-1 Semi-Private tasks ($0.02 per attempt) and 61.4% on ARC-AGI-2 Semi-Private tasks ($0.04 per attempt). It successfully passed the max reasoning tests on about half of the 120 public ARC-AGI-2 tasks.

Related event: DeepSeek V4 Flash Tops ARC-AGI Cost-Performance Frontier(4 posts)→

Original post →

More from Models

Models channel →