DeepSeek v4 Flash Sets New Cost/Accuracy SOTA on WeirdML

teortaxesTex · x · 2026-08-03

DeepSeek v4 Flash 0731 versions achieved 57.1% and 63.0% on WeirdML, setting a new cost/accuracy SOTA ahead of the recently repriced GPT 5.6 Luna.

Although the scores seem slightly below predictions, reviewers note the model often uses submissions just for data exploration rather than pure scoring. This suggests its true agentic capabilities are likely much higher, potentially reaching GLM 5.2 levels in the right harness.

Original post →

More from Models

Models channel →