Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost
Pawel Huryn's real-codebase benchmark found Google's cheapest agent model, Gemini 3.7 Flash, fixed fewer bugs than OpenAI's GPT-5.6 Luna while costing about 4.7 times more.
2026-08-25 ~ 2026-08-25 · 2 related posts
- Benchmark: Google's Gemini Flash Underperforms OpenAI's Luna in Bug Fixing and Cost — PawelHuryn · 2026-08-25
- Gemini 3.7 Flash loses to GPT-5.6 Luna in bug fix benchmark, costs 4.7x more — PawelHuryn · 2026-08-25