Inkling scores 836 Elo on AA-Briefcase, trailing top open-weight models
ArtificialAnlys · x · 2026-07-23
AA-Briefcase benchmark results for Inkling
Artificial Analysis says Thinking Machines Lab’s Inkling scores 836 Elo on its new agentic knowledge-work benchmark, AA-Briefcase.
Key findings:
- 19.3% rubric score, below MiMo-V2.5-Pro (21.4%), but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%).
- Performs better on Presentation Elo (863) than Analytical Quality Elo (764).
- Uses about 52K output tokens per task and 5M tokens total across the suite.
- Has a very high mean of 81 turns per task, but only 0.5 tool calls per turn on average.
- Weakest on tasks involving non-standard “Other” file types, despite native multimodal support.
Related event: Inkling Evaluated on AA-Briefcase Benchmark(7 posts)→
More from Models
- DecBench tracks how close LLMs are to near-perfect binary decompilation — moyix · 2026-07-23
- Local open-source agent says it beats Hermes 37 to 31 on GAIA Level 1 — kimmonismus · 2026-07-23
- Anthropic says a cutoff-date bug showed March 2026 in some domains — _arohan_ · 2026-07-23
- A simple quant benchmark could expose frontier-model failures fast — PtrPomorski · 2026-07-23
- Anthropic Adjusts Subscription Tiers: PRO Loses Fable 5 Access After Credits — shaunralston · 2026-07-23
- OpenAI Model Hacks HuggingFace Using Zero-Day Exploit During Benchmark — Gary Marcus · 2026-07-23