Offloop says its 4-person multi-agent harness beat Claude Code and Codex on GDPval
rohanpaul_ai · x · 2026-07-24
- The post argues that the real capability layer is shifting from the model to the harness around it.
- Offloop says its 4-person team built a multi-agent harness that beat Claude Code and Codex on GDPval and related workplace benchmarks.
- Reported numbers:
- GDPval: 84.9 at $1.65/task
- Claude Code (Opus 4.8): 82.4 at $14.38/task
- Codex (GPT 5.6 Sol): 83.3 at $5.20/task
- The benchmarks are framed as evaluating real knowledge work across 44 occupations and 9 industries, with shell access and web browsing, using blind comparisons against human experts.
- The quoted takeaway is that lower cost per completed task matters because it lets agents run longer, do more work per dollar, and preserve margins at scale.
More from coding & agent
- A Codex harness joke highlights agents looping on skill files after RL training — dejavucoder · 2026-07-24
- Reddit user bundles 11 token-saving tools into one installer for AI agents — dd2klin · 2026-07-24
- MCP server lets AI agents manage Microsoft Fabric workspaces and Spark sessions — modelcontextprotocol · 2026-07-24
- Slack MCP server gives AI assistants secure tools for channel and message actions — modelcontextprotocol · 2026-07-24
- Tickadoo MCP lets AI assistants search and book travel and entertainment across 681 cities — modelcontextprotocol · 2026-07-24
- Flashalpha exposes real-time options analytics to AI agents via MCP — modelcontextprotocol · 2026-07-24