SelfBench turns real GitHub PRs into evals: open-weight models cost more and do worse

ycombinator · x · 2026-10-06

SelfBench converts a repository's own merged pull requests into eval tasks to find the best coding model for your codebase. Tested on Next.js, Sentry, PostHog, Supabase, Vercel and more, open-weight models were usually pricier and less capable on real work.

Highlights:

The tool and harness are open source, so you can benchmark your own repo.

Original post →

More from coding & agent

coding & agent channel →