Inkling Evaluated on AA-Briefcase: Presentation Outperforms Analysis

Artificial Analysis released detailed evaluation data for Inkling on the new agentic knowledge work benchmark, AA-Briefcase. This benchmark tests models using real-world tasks that require processing massive amounts of files to deliver outputs like spreadsheets and presentations. Overall, Inkling achieved 836 Elo, lagging behind top open-source models, with a rubric score of 19.3%, placing it between MiMo-V2.5-Pro (21.4%) and DeepSeek V4 Flash max (18.7%).

Confirmed

In terms of capability distribution, Inkling shows a clear imbalance: its Presentation Elo stands at 863, significantly higher than its Analytical Quality Elo of 764. This indicates it is better at making final answers visually appealing and ready for delivery, but is relatively weaker in analytical structure and reasoning depth. Furthermore, despite claiming native multimodal support, Inkling loses the most points when handling non-standard "Other" file types (materials not in Excel, PowerPoint, PDF, or Word formats), making it its weakest area.

Regarding execution strategy and overhead, Inkling averages 81 interaction rounds per task (with a median of 49), suggesting it takes very long interaction chains for certain tasks. However, it only makes 0.5 tool calls per round on average, which is relatively low. In terms of resource consumption, the model consumes about 52,000 output tokens per task on average, requiring roughly 5 million tokens to run the entire benchmark.

2026-07-23 ~ 2026-07-23 · 7 related posts

Primary sources