Long-horizon agents still fail: best WeaveBench pass rate is just 41.2%

rohanpaul_ai · x · 2026-08-30

Rohan Paul reshares a discussion on long-horizon agent reliability: on WeaveBench's 114 hybrid GUI-CLI tasks, the best officially reported pass rate is just 41.2% — better models have not delivered long-horizon reliability.

Key points:

The quoted original is Chamath at Stanford AI Club: "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours." He predicts AI will hit a trough of disillusionment after the hype cycle.

Related event: Long-Horizon Agents Still Struggle as WeaveBench Top Pass Rate Hits Only 41.2%(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →