Dev builds private coding bench from his own git history; Flash Next Q2 tops it, Claude scores 8/12
fintip · reddit · 2026-09-22
Frustrated with benchmark gaming, a local-model enthusiast had Claude mine a decade of his own git history for real bugs and built a private coding benchmark that can't be benchmaxed: 100+ candidates, 12 tight core tests, 4-6 hours per run for local models.
Early results (12-test suite; locals via opencode, Gemini via agy-cli):
- Flash Next Q2 K XL leads at 11/12; slower per-token but far more efficient, matching 27B Q3 wall time while outperforming it
- The new Swift 27B tied for second at 10/11, fast and token-efficient
- 27B vanilla Q3 (unsloth) also 10/11
- Gemini 3.8 Flash (medium) 10/12
- Underdog Bonsai scored 8/10 (2 reruns pending after harness bugs)
- Corroborating recent "Claude got handicapped" chatter: Claude Code (open 5 + fable advisor) scored just 8/12, Haiku medium + fable also 8/12, Haiku low without fable 6/12
Caveat: the suite may over-index on Claude's weak spots since recent projects were written with Claude Code.
More from Models
- Xiaomi MiMo 2.6 Pro formalizes Li–Yorke chaos theorem in 6,000+ lines of verified Lean — bookwormengr · 2026-09-22
- Apollo, the first advanced LLM for Ancient Greek, aims to restore tattered papyri — nordicinst · 2026-09-22
- Two Minute Papers: Jev claims 200x AI speedup via System One models — but there's a catch — Two Minute Papers · 2026-09-22
- AI crowd discovers LLMs aren't always the cheapest, most effective tool — evilsocket · 2026-09-22
- cloneofsimo: academia badly underestimates the problems OpenAI's math agents are solving — cloneofsimo · 2026-09-22
- Founder finds asking the model to compare outputs restores drifting Astra quality — i_dg23 · 2026-09-22