Frontier Models Tested on Complex Agents: DeepSeek Wins on Cost Despite Inefficiency
rohanpaul_ai · x · 2026-08-06
A fascinating agent capability test tasked four frontier models—Qwen, DeepSeek, Kimi, and GPT-5.6—with building 3D Rubik's Cubes and chess boards, ultimately playing chess against Claude Opus 5 (all lost to Opus 5).
DeepSeek V4 Flash delivered a particularly interesting performance. Although it was highly inefficient at the token level—consuming 27.4 million tokens and requiring 5 attempts to build a working environment—it completed both complex tasks for only $0.557.
This suggests a shifting paradigm in evaluating agent efficiency: even if a model reasons inefficiently, it can remain economically viable and highly useful if inference is cheap enough to make retries virtually free.
More from coding & agent
- Winning with Claude Code Subagents: Cap Scope, Don't Just Unleash a Swarm — PrajwalTomar_ · 2026-08-06
- Less is More for AI Agents: Overloading Context Degrades Performance — jasonkneen · 2026-08-06
- Notion Opens Early Alpha for Custom Interactive Blocks Reading Workspace Data — ivanhzhao · 2026-08-06
- GraphARC Open-Source Framework: Building Controllable Investigation Agents with Local Qwen 8B — Desperate-Ad-9679 · 2026-08-06
- Dev Wraps Sports-Bet Simulator in MCP Server to Calculate True EV for LLMs — negative__ev · 2026-08-06
- 8-Year-Old Builds 32-Level Platformer Game in 4 Hours Using AI Studio — fofrAI · 2026-08-06