Inkling Evaluated on AA-Briefcase Benchmark

Artificial Analysis has released detailed evaluation data for Inkling on the new agentic knowledge work benchmark, AA-Briefcase. This benchmark tests model capabilities through real-world tasks involving massive file processing and the delivery of outputs like spreadsheets or presentations. Overall, Inkling scored 836 Elo, trailing top open-source models, with a rubric score of 19.3%—landing between MiMo-V2.5-Pro (21.4%) and DeepSeek V4 Flash max (18.7%).

Sub-Scores and Weaknesses

Inkling shows a highly uneven capability distribution: its presentation Elo stands at 863, significantly higher than its analysis quality Elo of 764. This indicates it is better at producing visually appealing, easy-to-deliver final answers, but relatively weaker in analytical structure and reasoning depth. Furthermore, despite claiming native multimodal support, Inkling lost the most points when handling non-standard "Other" file types (materials outside of Excel, PowerPoint, PDF, and Word), marking its weakest area.

Task Execution and Resource Consumption

Regarding execution strategy and overhead, Inkling averages 81 interaction turns per task (median of 49), indicating very long interaction chains for certain tasks. However, it averages only 0.5 tool calls per turn, which is relatively low. In terms of resource consumption, the model uses about 52,000 output tokens per task on average, meaning a full benchmark run requires approximately 5 million tokens.

2026-07-23 ~ 2026-07-23 · 7 related posts