Gemini 3.7 Flash loses to GPT-5.6 Luna in bug fix benchmark, costs 4.7x more

PawelHuryn · x · 2026-08-25

Benchmark results by Pawel Huryn reveal that Google's cheapest agent model, Gemini 3.7 Flash, underperforms compared to OpenAI's GPT-5.6 Luna. In a test of 105 planted bugs, Gemini 3.7 Flash fixed 22 bugs in 96 minutes at a cost of $8.43, while GPT-5.6 Luna fixed 33 bugs in 86 minutes for just $1.80. Gemini's performance also trailed behind Fable 5 and Grok 4.6, highlighting a significant price disadvantage.

Related event: Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost(2 posts)→

Original post →

More from Models

Models channel →