DocOps benchmark finds frontier agents still fail on long-horizon document tasks
Jiazhen Jiang · hf · 2026-07-23
DocOps introduces a verifiable benchmark for autonomous agents that operate on complex documents.
- The paper defines a hierarchical taxonomy that breaks document work into atomic actions and increasingly complex workflows.
- It evaluates representative closed- and open-source models across multiple agent harnesses.
- The authors find that even frontier systems still struggle with long-horizon, tightly coupled document tasks.
- Three recurring failure modes stand out: long-term state tracking collapse, shallow semantic verification, and destructive edits to structural metadata.
- The work is framed as a step toward more robust, non-destructive agents for real document-heavy workflows.
More from Research
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- Fast ViT shows strong ImageNet results; scaling runs needed next — ducha_aiki · 2026-09-11
- Loss Functions Are Scientific Assumptions: MSE Implies Gaussian Noise, Cross-Entropy Implies Bernoulli — bravo_abad · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- Researchers: LLMs under pressure invent new languages unreadable to humans — mikeflache · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11