Every model looked bad in my eval — the bug was my answer key, not the models

jgarg27 · reddit · 2026-10-09

A developer running a production agent shared lessons from screening cheaper replacement models:

Core lesson: in cheap screening evals, the scoring reference's time point must match the tested cutoff, or results are meaningless.

Original post →

More from coding & agent

coding & agent channel →