SelfBench turns real GitHub PRs into evals: open-weight models cost more and do worse
ycombinator · x · 2026-10-06
SelfBench converts a repository's own merged pull requests into eval tasks to find the best coding model for your codebase. Tested on Next.js, Sentry, PostHog, Supabase, Vercel and more, open-weight models were usually pricier and less capable on real work.
Highlights:
- Next.js: GPT-6.1 Sol most accurate at 84.2% ($0.91/task); also best value at 78.9% ($0.56)
- Sentry: Claude Sonnet 5.5 tops at 82.9% ($0.59/task)
- PostHog: Claude Opus 5.5 most accurate (69.4%) but $2.23/task
- vercel: Claude Opus 5.5 leads at 80.9%; budget pick Kimi K3 hits 78.7% ($0.96)
- Supabase: GPT-6.1 Sol dominates (77.8%, as low as $0.24/task)
The tool and harness are open source, so you can benchmark your own repo.
More from coding & agent
- Gradio says training your own models via a single ml-intern prompt is huge alpha — Gradio · 2026-10-06
- Most performance wins are under 5 lines of code — a 20% zstd fix case — DanielLockyer · 2026-10-06
- Developer laments that Claude Code and Codex do everything, leaving him out of the loop — zsakib_ · 2026-10-06
- Building agent skills from a structured wiki distilled from past experience — rseroter · 2026-10-06
- Using EvoX to draft bug reports: AI quietly turns "not shown" into "user skipped" — yawning42 · 2026-10-06
- Chunkr: open-source Rust chunking lib claims ~20x speedup over LangChain with full benchmarks — Ok_Cartographer5609 · 2026-10-06