Benchmark: Google's Gemini Flash Underperforms OpenAI's Luna in Bug Fixing and Cost

PawelHuryn · x · 2026-08-25

A benchmark test on real-world repositories shows that Google's cheapest agent model, Gemini 3.7 Flash, loses to OpenAI's GPT-5.6 Luna in both bug fixing performance and cost. Gemini 3.7 Flash fixed 22 out of 105 bugs in 96 minutes at $8.43, while GPT-5.6 Luna fixed 33 bugs in 86 minutes at just $1.80. The price gap is significant: Flash costs $3.75 per million output tokens, 3x more than Luna's $1.20, yet delivers inferior results.

Related event: Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost(2 posts)→

Original post →

More from coding & agent

coding & agent channel →