Bug-finding ability grows exponentially costlier across models, Paweł Huryn benchmark shows

garrytan · x · 2026-09-24

Paweł Huryn tested how well leading models find bugs planted in code and found that cost rises exponentially with performance (note the log scale), with ChatGPT used to superimpose a trend line: models above the line are the bargains. Paul Graham amplified the finding. The curve offers a quick sanity check on which model actually pays off for code review.

Original post →

More from coding & agent

coding & agent channel →