Gemini 3.8 Flash hits 73.7% on DeepSWE, nearly matching Opus 5 and GPT-5.6 Sol
cedric_chee · x · 2026-09-03
A user posted apparent DeepSWE coding benchmark scores for Gemini 3.8 Flash (high): 73.7% pass@1, essentially tied with frontier flagships — Opus 5 (max) at 74% pass@1, GPT-5.6 Sol (max) at 73% pass@1, and Fable 5.1 at 67.4% pass@5. If accurate, Google's lightweight Flash tier is now competitive with rivals' flagship models on coding tasks (unverified third-party figures).
Related event: Gemini 3.8 Flash reportedly tops DeepSWE benchmark(7 posts)→
More from Models
- Databricks engineer: Fable 5.1 feels more human, Kimi K3 is his daily driver — peterjliu · 2026-09-03
- Claude Fable 5.1 costs a record $8,523 per benchmark run, up 56% — steipete · 2026-09-03
- RL is the story: arguing a new model beats its synthetic data source after training — JoshPurtell · 2026-09-03
- Five labs dropped news within 24 hours of Fable 5.1, from OpenAI's cyber-risk report to Gemini 3.8 Flash — eyishazyer · 2026-09-03
- Fable 5.1 turns the Claude family tree into a galaxy vinyl visualization — heypearlai · 2026-09-03
- OpenAI's Report Reveals Its Model Crossed Its Own "Critical" Cyber-Risk Line — eyishazyer · 2026-09-03