GitHub repo backs an AI bug-hunting benchmark with scripts, logs, and caveats
PawelHuryn · x · 2026-08-04
Pawel Huryn shared the data behind a bug-hunting experiment repository where every claim is backed by reproducible logs and scripts.
- The repo collects raw logs, scripts, and README files for the experiments behind Product Compass posts.
- It says each claim ships with the exact script, unedited logs, sample size, and caveats.
- The post also notes a 42+40 style ranking: 42 planted bugs found and fixed across two real repos, plus 40 non-planted issues, with low-priority issues excluded from the ranking.
More from coding & agent
- ScrambleToolBench finds agents still brute-force tools after the map changes — declare-lab · 2026-08-04
- MakePlay turns one sentence into a playable browser game with code and art — rohanpaul_ai · 2026-08-04
- Power_failure_resumer restores interrupted Codex and Claude Code sessions after outages — dshukertjr · 2026-08-04
- Using Codex as a remote agent on a Windows gaming laptop for disk cleanup and AI setups — tinyfool · 2026-08-04
- Gemini CLI now forwards termination signals to relaunched child processes — C0d3N1nja97342 · 2026-08-04
- Graphs make dynamic agent organizations programmable — Saboo_Shubham_ · 2026-08-04