35B Beats 120B on Custom Coding Agent Benchmark, 95% vs 53% Success Rate

blaizedsouza · x · 2026-10-07

Engineer pauliusztin found a counterintuitive result while optimizing his coding agent: a 35B model hit a 95% success rate on his custom benchmark, while a model 3-4x larger at 120B managed only 53%.

He notes this is common when relying on generic benchmarks: high leaderboard scores say nothing about how a model, prompt, or technique performs on your own tasks — you have to measure on your own evals.

The post promotes lesson 7 of his open-source course Building a Coding Agent from Scratch (Agent Evals 101), covering:

Related event: Custom Agent Evals Show 35B Model Beating 120B Model(2 posts)→

Original post →

More from coding & agent

coding & agent channel →