Artificial Analysis Details AA-Briefcase: A Private Benchmark for Long-Horizon Agentic Workflows

ArtificialAnlys · x · 2026-08-13

Artificial Analysis has revealed more details about AA-Briefcase, its agentic knowledge work benchmark. The evaluation consists of four multi-week projects comprising thousands of input files and 91 tasks, covering realistic professional workflows in data science, product management, and corporate strategy.

Models are required to generate deliverables (such as spreadsheets, presentations, and memos) for each task, which are then graded against three dimensions: rubric checks, analytical quality, and presentation. A public fifth scenario has also been released on Hugging Face to illustrate the testing structure.

Original post →

More from Research

Research channel →