Multi-Model Bug Hunting Eval: Who Fixes Code Best

PawelHuryn · x · 2026-07-19

The author planted 45 bugs in their VS Code extension repo grok-build and tested five models under their respective native harnesses, using the same prompt and high effort conditions for bug finding and fixing.

The results show:

The author's recommended combo is: Fable 5 for judging/reviewing, and GPT-5.6 Sol for execution; if cost-effectiveness is the priority, run Grok 4.5 solo, and use Fable or Sol for backup when stuck. Ultimately, 27/45 bugs escaped all models.

Original post →

More from coding & agent

coding & agent channel →