Agent Evals 101: Benchmarking and Regression-Testing Agent Harnesses, 35B vs 120B

blaizedsouza · x · 2026-10-07

Paul Iusztin published "Agent Evals 101: Stop Guessing, Start Measuring," a guide to building benchmarks and regression tests for any agent harness.

While optimizing his coding agent on a custom benchmark, he compared a 35B vs. a 120B model—the larger one, despite being 3-4x bigger, didn't automatically win, showing intuition alone can't guide agent optimization.

The core message: measure instead of guessing, and use reproducible eval pipelines to regression-test agent changes and catch silent degradations.

Related event: Custom Agent Evals Show 35B Model Beating 120B Model(2 posts)→

Original post →

More from coding & agent

coding & agent channel →