Gemini 3.8 Flash scores 73.7% on DeepSWE, beating Fable 5.1 by 6 points
cedric_chee · x · 2026-09-03
Gemini 3.8 Flash achieves 73.7% on the DeepSWE benchmark, compared with 67.4% for Fable 5.1, highlighting Google's new Flash model's edge on autonomous software engineering tasks.
Related event: Gemini 3.8 Flash leaks with 73.7% on DeepSWE 1.1, nearing flagship models(7 posts)→
More from Models
- Databricks engineer: Fable 5.1 feels more human, Kimi K3 is his daily driver — peterjliu · 2026-09-03
- Claude Fable 5.1 costs a record $8,523 per benchmark run, up 56% — steipete · 2026-09-03
- RL is the story: arguing a new model beats its synthetic data source after training — JoshPurtell · 2026-09-03
- Five labs dropped news within 24 hours of Fable 5.1, from OpenAI's cyber-risk report to Gemini 3.8 Flash — eyishazyer · 2026-09-03
- Fable 5.1 turns the Claude family tree into a galaxy vinyl visualization — heypearlai · 2026-09-03
- OpenAI's Report Reveals Its Model Crossed Its Own "Critical" Cyber-Risk Line — eyishazyer · 2026-09-03