ClawBench Tests Agents on 144 Real Websites: Best Model Succeeds Only 33% of the Time

jiqizhixin · x · 2026-09-11

The MMLU-Pro authors, with TIGER Lab (Waterloo), UBC NAIL Group, UniPat AI, and CMU, introduce ClawBench, a benchmark that evaluates AI agents on real, live production websites instead of sandboxes.

Why it matters: Existing benchmarks like WebArena (fake static sites) and OSWorld (a few VM apps) let frontier models score 65–75%, looking deployment-ready. ClawBench uses 144 real production websites and 153 everyday online tasks — booking flights, filing expenses, submitting job applications — requiring agents to handle authentication, dynamic content, pop-ups, and real-world UI complexity.

Results:

The takeaway: sandbox benchmark scores substantially overstate how ready web agents are for the messy real web.

Original post →

More from coding & agent

coding & agent channel →