Open-Sourced AI Agent Experiment Repo: Bug Hunting and Model Benchmarks
PawelHuryn · x · 2026-08-01
Pawel Huryn of Product Compass has open-sourced a GitHub repository containing raw data and scripts from AI model and agent experiments featured in his blog.
The repo emphasizes full reproducibility, offering exact scripts, unedited logs, and READMEs detailing methodology, sample sizes, and caveats. Current experiments include:
- Bug Hunt Bench: Testing 105 bugs across two real-world repos.
- Model & Harness Comparisons: Evaluating frontier vs. open-source models, and managed vs. local agents.
- Specific Tests: Claude Code prompt shrinking, Kimi K3 day-one tests, and chess playing evaluations.
More from coding & agent
- Omnigent Tested: Multi-Agent Coding and Cross-Review Workflow Boosts Efficiency — myth-buster9999 · 2026-08-01
- Full Automation: Letting Agents Handle Model Evaluation End-to-End — AlchainHust · 2026-08-01
- Cut Your $200/Month AI Bill: 10 Steps to Local AI Deployment — jason_mayes · 2026-08-01
- Developer Accidentally Leaves Codex Agent Running for 5 Days — vxnuaj · 2026-08-01
- GeoAI: Open-Source Python Library Integrates Deep Learning with Geospatial Data — tom_doerr · 2026-08-01
- DataTalks.Club Webinar: Building an End-to-End Automated RAG FAQ Assistant — Al_Grigor · 2026-08-01