DocOps benchmark finds frontier agents still fail on long-horizon document tasks
Jiazhen Jiang · hf · 2026-07-23
DocOps introduces a verifiable benchmark for autonomous agents that operate on complex documents.
- The paper defines a hierarchical taxonomy that breaks document work into atomic actions and increasingly complex workflows.
- It evaluates representative closed- and open-source models across multiple agent harnesses.
- The authors find that even frontier systems still struggle with long-horizon, tightly coupled document tasks.
- Three recurring failure modes stand out: long-term state tracking collapse, shallow semantic verification, and destructive edits to structural metadata.
- The work is framed as a step toward more robust, non-destructive agents for real document-heavy workflows.
More from Research
- New ASCIITermDraw benchmark says top VLMs still miss simple text diagrams — East-Muffin-6472 · 2026-07-23
- Stanford’s vine-like soft robot grows from the tip to reach trapped people — lukas_m_ziegler · 2026-07-23
- First CAR-T Cell Therapy Approved for Solid Tumors in Gastric Cancer — Dr_Singularity · 2026-07-23
- A production multi-agent team says deterministic orchestration works better than deterministic LLMs — njanChe1 · 2026-07-23
- LxMLS 2026 shares a public video-lecture collection from Lisbon Machine Learning School — caglar_ee · 2026-07-23
- WSD Schedule vs. Early Stopping: Open Speech Model Training Post-mortem — irombie · 2026-07-23