Android Bench Updates Model Leaderboard
Ars Technica AI · rss · 2026-07-09
Google updated Android Bench, its benchmark for evaluating Android app development capabilities, adding several new models and evaluation dimensions.
This update incorporates cost and efficiency metrics alongside open-weight models, helping developers comprehensively compare LLM performance on Android tasks. Google also invites developers to run tests and provide feedback to shape the benchmark's future.
Newly added models include: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, Qwen 3.7 Max. The article concludes that Gemini still trails some competing models on this leaderboard.
More from Models
- giffmana: the env being used in training is part of the point — giffmana · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11