Every model looked bad in my eval — the bug was my answer key, not the models
jgarg27 · reddit · 2026-10-09
A developer running a production agent shared lessons from screening cheaper replacement models:
- Structural tests worked well: does the model call its tools, is the output schema-valid, is the first step sensible? Cheap, fast, and clearly separating models. One gotcha: a seemingly weak model was just using a different tool-call format.
- Judgment tests failed: cut each run midway, ask the model to predict the conclusion, score against the agent's final answer — every model, including the production one, looked terrible.
- Diagnosis: testing the production model against its own midpoint decisions matched perfectly, so the harness was fine. The bug was the answer key — it came from the end of the run, but the test stopped earlier. "Grading a midpoint against a finish line."
- After fixing the cutoff, most cheaper models matched the production model; only the weakest didn't.
- Cost caveats: cheapest open-weight models were slow on managed endpoints, one hit rate limits at low parallelism — cheaper per call ≠ cheaper per decision.
Core lesson: in cheap screening evals, the scoring reference's time point must match the tested cutoff, or results are meaningless.
More from coding & agent
- TinyJoin v0.6 now beats SQLite and PGlite on 20 benchmarks, stays smaller — Vjeux · 2026-10-09
- Dev Ports Muse Linux Gadget SDK to .NET, Gets Edge Running on Arduino — unixterminal · 2026-10-09
- Eazo bets personal agents should become reusable apps, not chat windows — alifcoder · 2026-10-09
- Qoni Launches Personal Agent Infrastructure: Identity, Action and Memory as Plug-In Layers — alifcoder · 2026-10-09
- Vibe-founder lets Codex run a startup overnight, signs 10 business customers in 24h — CtrlAltDwayne · 2026-10-09
- boat spins up 250 agent sandbox VMs for 10 cents: full Ubuntu boxes at $20/mo — RexDouglass · 2026-10-09