Dev builds private coding bench from his own git history; Flash Next Q2 tops it, Claude scores 8/12

fintip · reddit · 2026-09-22

Frustrated with benchmark gaming, a local-model enthusiast had Claude mine a decade of his own git history for real bugs and built a private coding benchmark that can't be benchmaxed: 100+ candidates, 12 tight core tests, 4-6 hours per run for local models.

Early results (12-test suite; locals via opencode, Gemini via agy-cli):

Caveat: the suite may over-index on Claude's weak spots since recent projects were written with Claude Code.

Original post →

More from Models

Models channel →