ClawWork benchmark gives AI agents $10 and ranks them by simulated wages and token costs
aigclink · x · 2026-07-26
A post highlights ClawWork, an agent benchmark that simulates an AI employee earning wages, spending tokens, and possibly going bankrupt.
How it works
- The agent starts with $10.
- It is given GDPVal tasks: 220 real-world职业 tasks across 44 occupations.
- Every token spent is deducted at real API rates.
- A submitted result is scored by GPT using job-specific rubrics.
- Simulated income = quality score × BLS hourly wage × time worked.
What the results show
- Overall quality is relatively low: the top agent, ATIC + Qwen3.5-Plus, averages 61.6%.
- Most models cluster around 36%–43%.
- The highest-quality agent, ATIC-DEEPSEEK at 66.8%, only ranks 5th in earnings.
- The earnings leaderboard is led by ATIC + Qwen3.5-Plus.
The post argues the system’s architecture is worth studying as a template for future benchmarks of AI workers.
Related event: ClawWork: AI Agents Go to Work with Real API Costs(2 posts)→
More from coding & agent
- RTK is a Rust CLI proxy that cuts AI agent shell output by up to 90% — OpenAIDevs · 2026-07-26
- Daniel Lockyer recommends Moshi and herdrdev for remote agent coding and multiplexing — DanielLockyer · 2026-07-26
- ChatGPT Voice reportedly completed a Meta Business Suite flow hands-free — OpenAIDevs · 2026-07-26
- Open-source project targets the 90% of terminal output that coding agents waste context on — thetripathi58 · 2026-07-26
- RTK: Open-Source CLI Proxy Compresses Terminal Noise for AI Coding Agents — thetripathi58 · 2026-07-26
- Uncle Bob says agents need constant monitoring, context resets, and deterministic guardrails — NathanWilbanks_ · 2026-07-26