Just 3 Weeks Later: Gemini 3.7 Flash Shows Major Gains in Coding and Document Processing
Ars Technica AI · rss · 2026-08-14
Google has rapidly introduced Gemini 3.7 Flash merely three weeks after the release of 3.6 Flash. Positioned as a "workhorse" model, it achieves significant improvements in coding and agentic performance, coupled with a lower introductory price.
- Coding Performance: The FrontierCode 1.1 test score jumped from 34.4% to 43.6%, DeepSWE v1.1 went from 49% to 65.3%, and the WebDev Arena score rose to 1,588.
- Complex Docs & Automation: The GDP.pdf benchmark, measuring complex document processing, increased from 22% to 34%. AutomationBench, evaluating common business workflows, doubled its score from 17% to 30.4%.
More from Models
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Sakana AI translation outperforms Google and DeepL in Japanese-English benchmarks — SakanaAILabs · 2026-08-24
- Developer haider makes his own LLM tier list after disagreeing with theo's rankings — haider1 · 2026-08-24
- Mystery OxAlpha Beats Claude; Alibaba Raises $10B for AI — 创业邦 · 2026-08-24
- OpenAI and Google cut LLM prices; mystery OxAlpha model beats Claude on DeepSWE — 创业邦 · 2026-08-24
- AI News Digest: DeepSeek Weekend Discounts, GPT-5.6 Sol Price Cut, Alibaba's $10B AI Raise — APPSO · 2026-08-24