Self-bench: Open-source tool to auto-build private coding-agent benchmarks from your repo
ricklamers · x · 2026-08-15
self-bench is an open-source package that builds private coding-agent benchmarks from work already completed in your repository, such as merged PRs. It reconstructs tasks from pre-change commits, generates hidden tests and reference solutions, and exports to Harbor format for comparing models on your real codebase.
More from coding & agent
- OpenRouter launches Ori DeepSeek Harness for easy deployment — gaganghotra_ · 2026-08-15
- OpenRouter launches Ori Prime Agent with access to 500+ models — samsja19 · 2026-08-15
- AI Engineer World's Fair 2026: Computer Use Track Agenda — proceduralia · 2026-08-15
- Hermes Agent now supports exporting entire agent configuration to a single file — Saboo_Shubham_ · 2026-08-15
- Discussion: Boundaries and risks of write access for AI Agents in production — justinotherflow · 2026-08-15
- PIRT: Run a Linux Agent Directly on Your Android Phone — Glad_Raspberry_6795 · 2026-08-15