Gemini 3.8 Flash edges out Astra on DeepSWE: 73.8% vs 73.3%
jon_barron · x · 2026-09-04
Retweeting @theo's comparison: Gemini 3.8 Flash scores 73.8% on the DeepSWE coding benchmark, narrowly beating Astra's 73.3% by half a point.
More from Models
- Deep Learning Weekly #471: Claude Fable 5.1 launch, production-parity LLM evals, alignment paper — dl_weekly · 2026-09-05
- OpenAI confirms Astra counts toward normal plan usage, users can allocate 100% of quota — soumitrashukla9 · 2026-09-05
- RareBench eval: Claude Fable 5.1 tops rare-disease diagnosis while Nemotron 3 Ultra scores 0% — danielmckinn0n · 2026-09-05
- GPT-6 Astra Beats 5.6 Sol Pro (Max) on FrontierMath T4; Open Models Seen 18 Months Behind — inductionheads · 2026-09-05
- TheZvi breaks down the Claude Fable 5.1 system card: 200+ pages of safety evals — TheZvi · 2026-09-05
- COLM paper: legible chain-of-thought steps aren't necessarily important — LauraRuis · 2026-09-05