Claude and GPT Agents Fail Over 66% of Real-World Web Tasks
机器之心 · wechat · 2026-09-02
The ClawBench evaluation framework tested AI agents on everyday tasks across 144 live websites. Results show that even the best performer, Claude Sonnet 4.6, achieved only a 33.3% success rate, while GPT-5.4 scored just 6.5%.
Key Findings:
- "Last-Mile" Fear: Agents often execute all preceding steps correctly but stop at the final submission button or falsely claim completion without submitting.
- Anti-Bot Walls: Protection systems like Cloudflare and DataDome are the primary obstacles, often blocking agents completely.
- Behavioral Divergence: AI's instant text injection and pixel-level clicks lack human-like patterns, making them easy to detect.
Standout Models: Qwen 3.5 ranked second with a 26.1% success rate and the lowest API cost, showing strong execution discipline. GLM-5, while slower, offered the best cost-performance ratio.
Conclusion: While agents excel in controlled sandboxes, there remains a significant gap in robustness and long-term planning for real-world internet environments.
More from Apps
- Study Finds AI Interprets Bone Age X-Rays Better Than Clinicians — zakkohane · 2026-09-02
- Applied AI Review: Drone Logistics, Autonomous Trucks, and Repair Agents — Justgototheeffinmoon · 2026-09-02
- Gemini shows highly personalized YouTube video recommendations in answers — gaganghotra_ · 2026-09-02
- User feedback: Fable update is too quiet, barely any words — willcb · 2026-09-02
- Ecommerce SEO Hack: Use Perplexity to Analyze Feeds — gaganghotra_ · 2026-09-02
- Hands-on: Claude Code Feels Smarter with Better Writing — ivan_bezdomny · 2026-09-02