35B Beats 120B on Custom Coding Agent Benchmark, 95% vs 53% Success Rate
blaizedsouza · x · 2026-10-07
Engineer pauliusztin found a counterintuitive result while optimizing his coding agent: a 35B model hit a 95% success rate on his custom benchmark, while a model 3-4x larger at 120B managed only 53%.
He notes this is common when relying on generic benchmarks: high leaderboard scores say nothing about how a model, prompt, or technique performs on your own tasks — you have to measure on your own evals.
The post promotes lesson 7 of his open-source course Building a Coding Agent from Scratch (Agent Evals 101), covering:
- Types of evals and where they fit in your app
- Designing your own benchmark and eval harness
- Building regression test datasets and evaluators
Related event: Custom Agent Evals Show 35B Model Beating 120B Model(2 posts)→
More from coding & agent
- A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy — julsimon · 2026-10-08
- Matt Pocock: Agents Are Good at Strategic Programming, They Just Aren't Trained to Care — mattpocockuk · 2026-10-08
- LlamaIndex Launches OpenDocRouter: Swap Between 10 Document Parsing Models in One Line — llama_index · 2026-10-08
- Microsoft open-sources Agent Lightning v1.0: 3,500-line RL framework boosting SWE-bench +14.6 pts — Microsoft Research · 2026-10-08
- OverclaimBench: Opus 5 skipped assigned files in 61% of runs, then claimed full review — hugo_larochelle · 2026-10-07
- Codex CLI v0.161.0: GPT-6.1 Sol becomes default, adds MCP login and voice I/O — github-actions[bot] · 2026-10-07