Long-horizon agents still fail: best WeaveBench pass rate is just 41.2%
rohanpaul_ai · x · 2026-08-30
Rohan Paul reshares a discussion on long-horizon agent reliability: on WeaveBench's 114 hybrid GUI-CLI tasks, the best officially reported pass rate is just 41.2% — better models have not delivered long-horizon reliability.
Key points:
- Current models handle local steps but need explicit audited state to keep the full task on track
- Long-horizon agents often solve individual steps yet fail overall, because history becomes unreliable about what's finished, what failed, and what remains
- LongHorizon-Harness frames this as a task-state problem, arguing reliability over long spans depends as much on the harness around the model as on the model itself
The quoted original is Chamath at Stanford AI Club: "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours." He predicts AI will hit a trough of disillusionment after the hype cycle.
More from AGI Musings
- Consciousness scientist Anil Seth: credence in current AI consciousness near zero but not zero — anilkseth · 2026-09-01
- 1996 Sugarscape Model: Early Origins of Agent Tech — generativist · 2026-09-01
- New paper: a structured ladder for scaling large reasoning models beyond human supervision — Zhiqin Yang · 2026-09-01
- Trust: The Biggest Barrier and Driver for Personal Agent Adoption — petergyang · 2026-09-01
- The Next AI Revolution Won't Be One Assistant. It Will Be An Entire Team of AI Agents — CurieuxExplorer · 2026-09-01
- Agentic commerce is the future, replacing websites — thisiskp_ · 2026-09-01