Gemini 3.8 Flash scores 1213 Elo on AA-Briefcase, up 79 points over predecessor
ArtificialAnlys · x · 2026-09-02
On AA-Briefcase, Artificial Analysis' flagship agentic knowledge work eval built on a fully private dataset of realistic tasks, Gemini 3.8 Flash scores an Elo of 1213, up 79 points from Gemini 3.7 Flash (1134). The gain comes mostly from higher rubric scores and Analytical Quality rather than improved Presentation Quality.
Related event: Gemini 3.8 Flash tops price-performance Pareto frontier(9 posts)→
More from Models
- Google launches Gemini 3.8 Flash Cyber security model alongside Fairwind Program for defenders — GoogleAI · 2026-09-03
- Google launches Gemini 3.8 Flash Cyber, helping Chrome team produce 2.6x more correct patches — GoogleAI · 2026-09-03
- Matt Shumer: slow AI releases aren't a wall — safety clearance is the bottleneck, wave of frontier models imminent — mattshumer_ · 2026-09-03
- Fable 5.1 draws strongest first impressions yet, but burns through rate limits far faster — JosephJacks_ · 2026-09-03
- Unverified claim: Gemini 3.8 Flash tops DeepSWE v1.1 coding benchmark — vedantmisra · 2026-09-03
- Early Gemini 3.8 testing suggests Gemma 4.5 will be very good — andrew_n_carr · 2026-09-03