Benchmark: Google's Gemini Flash Underperforms OpenAI's Luna in Bug Fixing and Cost
PawelHuryn · x · 2026-08-25
A benchmark test on real-world repositories shows that Google's cheapest agent model, Gemini 3.7 Flash, loses to OpenAI's GPT-5.6 Luna in both bug fixing performance and cost. Gemini 3.7 Flash fixed 22 out of 105 bugs in 96 minutes at $8.43, while GPT-5.6 Luna fixed 33 bugs in 86 minutes at just $1.80. The price gap is significant: Flash costs $3.75 per million output tokens, 3x more than Luna's $1.20, yet delivers inferior results.
Related event: Benchmark: Gemini Flash Loses to GPT-5.6 Luna on Bug Fixing at 4.7x Cost(2 posts)→
More from coding & agent
- VecturaKit: Swift-based on-device vector database with MLX acceleration — rudrank · 2026-08-25
- Building a Multi-Agent Personal Assistant Workflow with Claude — michael_k18 · 2026-08-25
- Browser Use CLI reportedly boosts Hermes agent performance 10x — intellectronica · 2026-08-25
- session-migrate: move coding agent sessions across Claude Code, Codex in one command — xhluca · 2026-08-25
- SUCCESSOR Ω: Neural-Symbolic System Generates and Evolves Executable World Programs — Ghost_Pilot_MD · 2026-08-25
- AI Fact-Checker Audit: 1 in 18 Citations Were Fabricated — jonathancheckwise · 2026-08-25