Frontier Models Tested on Complex Agents: DeepSeek Wins on Cost Despite Inefficiency

rohanpaul_ai · x · 2026-08-06

A fascinating agent capability test tasked four frontier models—Qwen, DeepSeek, Kimi, and GPT-5.6—with building 3D Rubik's Cubes and chess boards, ultimately playing chess against Claude Opus 5 (all lost to Opus 5).

DeepSeek V4 Flash delivered a particularly interesting performance. Although it was highly inefficient at the token level—consuming 27.4 million tokens and requiring 5 attempts to build a working environment—it completed both complex tasks for only $0.557.

This suggests a shifting paradigm in evaluating agent efficiency: even if a model reasons inefficiently, it can remain economically viable and highly useful if inference is cheap enough to make retries virtually free.

Original post →

More from coding & agent

coding & agent channel →