Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost

Pawel Huryn's real-codebase benchmark found Google's cheapest agent model, Gemini 3.7 Flash, fixed fewer bugs than OpenAI's GPT-5.6 Luna while costing about 4.7 times more.

2026-08-25 ~ 2026-08-25 · 2 related posts