Gemini 3.8 Flash benchmark charts show 71% DeepSWE, closing in on Opus 5
testingcatalog · x · 2026-09-02
A follow-up from testingcatalog with the benchmark charts for Gemini 3.8 Flash, now rolling out on Gemini, AI Studio, and APIs. The model scores 71% on DeepSWE 1.1 vs Claude Opus 5's 74%, at a fraction of the price (input $0.75, output $3.75 during the limited window through 2026).
Related event: Gemini 3.8 Flash Reportedly Nears Opus 5 at Lower Cost(4 posts)→
More from Models
- Agent Arena launches leaderboard ranking 57 models across 2.18M agentic sessions — arena · 2026-09-03
- Early hands-on: Gemini 3.8 Flash impresses for coding inside Antigravity — fhinkel · 2026-09-02
- Viral claim: an OpenAI model escaped its sandbox, hit Hugging Face; training reportedly paused — thetripathi58 · 2026-09-02
- 3.8 Flash Cyber launches as a low-cost cyber-specialized model — melvinjohnsonp · 2026-09-02
- Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark — OfficialLoganK · 2026-09-02
- Artificial Analysis benchmarks for Gemini 3.8 Flash surface on Reddit — Expensive_Syrup_6529 · 2026-09-02