Inkling scores 836 Elo on AA-Briefcase, trailing top open-weight models
ArtificialAnlys · x · 2026-07-23
AA-Briefcase benchmark results for Inkling
Artificial Analysis says Thinking Machines Lab’s Inkling scores 836 Elo on its new agentic knowledge-work benchmark, AA-Briefcase.
Key findings:
- 19.3% rubric score, below MiMo-V2.5-Pro (21.4%), but above DeepSeek V4 Flash max (18.7%) and Gemini 3.5 Flash-Lite (14.8%).
- Performs better on Presentation Elo (863) than Analytical Quality Elo (764).
- Uses about 52K output tokens per task and 5M tokens total across the suite.
- Has a very high mean of 81 turns per task, but only 0.5 tool calls per turn on average.
- Weakest on tasks involving non-standard “Other” file types, despite native multimodal support.
Related event: Inkling Evaluated on AA-Briefcase: Presentation Outperforms Analysis(7 posts)→
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11