Agent Evals 101: Benchmarking and Regression-Testing Agent Harnesses, 35B vs 120B
blaizedsouza · x · 2026-10-07
Paul Iusztin published "Agent Evals 101: Stop Guessing, Start Measuring," a guide to building benchmarks and regression tests for any agent harness.
While optimizing his coding agent on a custom benchmark, he compared a 35B vs. a 120B model—the larger one, despite being 3-4x bigger, didn't automatically win, showing intuition alone can't guide agent optimization.
The core message: measure instead of guessing, and use reproducible eval pipelines to regression-test agent changes and catch silent degradations.
Related event: Custom Agent Evals Show 35B Model Beating 120B Model(2 posts)→
More from coding & agent
- Codex Can Now Control Electron Apps and Test, Debug Them Directly — JeremyNguyenPhD · 2026-10-07
- AI Security Institute open-sources Transect, turning agent transcripts into interactive timelines — joshua_saxe · 2026-10-07
- Brian Balfour: building an agentic growth mechanic when AI makes growth a search problem — morganb · 2026-10-07
- "Agent says it's done" isn't ready: passing tests only cover a fraction of ship-readiness — Glittering-Glass6135 · 2026-10-07
- A changelog prompt keeps your Grok Bot updated on its own improvements in long-running jobs — RachelVT42 · 2026-10-07
- Video background replacement with H3 Inpainting + LTX Alpha Matte LoRA, plus a diffusers workflow — linoy_tsaban · 2026-10-07