Bug-finding ability grows exponentially costlier across models, Paweł Huryn benchmark shows
garrytan · x · 2026-09-24
Paweł Huryn tested how well leading models find bugs planted in code and found that cost rises exponentially with performance (note the log scale), with ChatGPT used to superimpose a trend line: models above the line are the bargains. Paul Graham amplified the finding. The curve offers a quick sanity check on which model actually pays off for code review.
More from coding & agent
- AI-coded 3D Peach Blossom Spring web scene open-sourced, cost 10-15% of a week's Max quota — dotey · 2026-09-24
- Qwen3.8-Flash free in Qoder through Sept 30, no Credits for free accounts — SimplyAnnisa · 2026-09-24
- Opus 5.5 keeps trying to run rm -f; users must repeatedly beg it not to — gandamu_ml · 2026-09-24
- One-year bet: shopping, research and seller agents trading over MPP with stablecoins — schwentker · 2026-09-24
- AWS on agent spending limits: demo wallet held $1M but agent session budget was 15 cents — schwentker · 2026-09-24
- Stripe launches Directory in preview, letting AI agents find providers via one CLI command — schwentker · 2026-09-24