Long-Horizon Agents Still Struggle as WeaveBench Top Pass Rate Hits Only 41.2%
Long-horizon agents remain unreliable, with the best reported pass rate on WeaveBench's 114 mixed GUI-CLI tasks at just 41.2%. A new framework, LongHorizon-Harness, addresses this by externalizing task state and adding audit loops, tripling scores on OSWorld.
2026-08-30 ~ 2026-08-30 · 2 related posts
- LongHorizon-Harness: external task state + audit loop triples OSWorld agent scores — rohanpaul_ai · 2026-08-30
- Long-horizon agents still fail: best WeaveBench pass rate is just 41.2% — rohanpaul_ai · 2026-08-30