Long-Horizon Agents Still Struggle as WeaveBench Top Pass Rate Hits Only 41.2%

Long-horizon agents remain unreliable, with the best reported pass rate on WeaveBench's 114 mixed GUI-CLI tasks at just 41.2%. A new framework, LongHorizon-Harness, addresses this by externalizing task state and adding audit loops, tripling scores on OSWorld.

2026-08-30 ~ 2026-08-30 · 2 related posts