Gemini 3.7 Flash loses to GPT-5.6 Luna in bug fix benchmark, costs 4.7x more
PawelHuryn · x · 2026-08-25
Benchmark results by Pawel Huryn reveal that Google's cheapest agent model, Gemini 3.7 Flash, underperforms compared to OpenAI's GPT-5.6 Luna. In a test of 105 planted bugs, Gemini 3.7 Flash fixed 22 bugs in 96 minutes at a cost of $8.43, while GPT-5.6 Luna fixed 33 bugs in 86 minutes for just $1.80. Gemini's performance also trailed behind Fable 5 and Grok 4.6, highlighting a significant price disadvantage.
Related event: Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost(2 posts)→
More from Models
- Grok generates unprompted image push to boost app engagement — StefanoGogioso · 2026-08-25
- Security Researcher Waits a Month for Claude Cyber Trusted Access Approval — nptacek · 2026-08-25
- Codex overage allowance slashed to ~1%; exploit value > disclosure bounty — nptacek · 2026-08-25
- Developer complains about Ox Alpha's slow inference: 127 mins for 10 min task — altryne · 2026-08-25
- a16z partner blown away by access to unreleased AI model — AccBalanced · 2026-08-25
- Chinese LLMs 4-5 Months Behind US; ECI 155 May Be Reliability Threshold — Jsevillamol · 2026-08-25