Gemini 3.8 Flash ties for top of DeepSWE at 74%, but burns 1.34x more output tokens
haider1 · x · 2026-09-03
- Top of DeepSWE: Gemini 3.8 Flash now ties at the top of DeepSWE, a benchmark for long-horizon software engineering tasks, scoring 74%, up from Gemini 3.7 Flash's 65%.
- Token efficiency takes a hit: it outputs about 143k tokens per task vs 107k for 3.7 Flash — 1.34x more.
- Cross-model comparison: the author notes it scores 59 on the AA Intelligence Index, a big jump, yet uses more output tokens per task than the GPT-5.6 family, Opus 5, Fable 5, and even Fable 5.1 — intelligence improved, but token efficiency looks rough.
Related event: Gemini 3.8 Flash benchmarks and rollout leak ahead of launch(42 posts)→
More from Models
- Leak: Astra won't be the best model of the year; a 'monster' is slated for end of year — ChrisGPT · 2026-09-03
- Startup Mostik bridges AI models via their weights, tops ARC-AGI 3 at 1/20 the cost — nordicinst · 2026-09-03
- Anthropic launches browser tool to detect Claude-made files via C2PA content credentials — btibor91 · 2026-09-03
- ByteDance's looped language models match 12B rivals at 1.4B size, with Bengio as co-author — peterjliu · 2026-09-03
- Anthropic Weakened Safety Filters, Signed an AI Cyberattack Warning Letter, Then Shipped Mythos 5.1 Anyway — AgentBlackVeil · 2026-09-03
- Marin 535B A23B Frontier-Scale Training Run Is Fully Livestreamed — Sentdex · 2026-09-03