New OfficeQA Pro V2 Benchmark: Frontier AI Agents Average Just 26% Accuracy
jefrankle · x · 2026-08-07
Databricks has introduced OfficeQA Pro V2, a new benchmark designed to test AI agents on complex document processing.
- Dataset: Built on a corpus of 120,000 PDFs provided by the U.S. Treasury.
- Performance: Frontier AI agents average only 26% accuracy, highlighting significant challenges in real-world document comprehension.
Related event: Databricks Launches OfficeQA Pro V2 Benchmark(2 posts)→
More from Research
- AI-Designed Viruses Spark Nordic Biosecurity Debate — nordicinst · 2026-08-07
- Extracting Robotic Action Signals from Egocentric Videos Using Only Open-Source Models — gui_penedo · 2026-08-07
- Testing 13 Agent Search APIs: Hidden Token Costs Vary by 67x — Patient-Injury-1327 · 2026-08-07
- EgoHumanoid Framework: Egocentric Human Demos Boost Robot Generalization by 51% — micoolcho · 2026-08-07
- Cohere Partners with Meta and DeepMind for ML Summer School — Cohere · 2026-08-07
- Cooperative Identities in Model Instances Can Emerge Without RL — jankulveit · 2026-08-07