ClawBench Tests Agents on 144 Real Websites: Best Model Succeeds Only 33% of the Time
jiqizhixin · x · 2026-09-11
The MMLU-Pro authors, with TIGER Lab (Waterloo), UBC NAIL Group, UniPat AI, and CMU, introduce ClawBench, a benchmark that evaluates AI agents on real, live production websites instead of sandboxes.
Why it matters: Existing benchmarks like WebArena (fake static sites) and OSWorld (a few VM apps) let frontier models score 65–75%, looking deployment-ready. ClawBench uses 144 real production websites and 153 everyday online tasks — booking flights, filing expenses, submitting job applications — requiring agents to handle authentication, dynamic content, pop-ups, and real-world UI complexity.
Results:
- Best model Claude Sonnet 4.6 succeeds on just 33% of tasks
- GPT-5.4 completes only 10 of 153 tasks (6.5%)
- 68 of 153 tasks (44.4%) are solved by zero models
The takeaway: sandbox benchmark scores substantially overstate how ready web agents are for the messy real web.
More from coding & agent
- Muse agent browser blocks pasting, frustrating login flows; 1Password integration requested — altryne · 2026-09-11
- Greg Isenberg: GPT-6 Astra unlocks physical product startups that needed $2M two years ago — Rasmic · 2026-09-11
- Stress test: running atelier nested inside itself works and stays fast — lucasmeijer · 2026-09-11
- antirez Is Working on DeepSeek V4.1 Support for His ds4 Editor — backyard_tractorbeam · 2026-09-11
- The hard part of agents was never the model — it's hours-long reliability, says dev on OpenAI's Agents API — shaunralston · 2026-09-11
- MacBook lid angle sensor powers a folding shader app, now open-sourced — jh3yy · 2026-09-11