Claude and GPT Agents Fail Over 66% of Real-World Web Tasks

机器之心 · wechat · 2026-09-02

The ClawBench evaluation framework tested AI agents on everyday tasks across 144 live websites. Results show that even the best performer, Claude Sonnet 4.6, achieved only a 33.3% success rate, while GPT-5.4 scored just 6.5%.

Key Findings:

Standout Models: Qwen 3.5 ranked second with a 26.1% success rate and the lowest API cost, showing strong execution discipline. GLM-5, while slower, offered the best cost-performance ratio.

Conclusion: While agents excel in controlled sandboxes, there remains a significant gap in robustness and long-term planning for real-world internet environments.

Original post →

More from Apps

Apps channel →