Agent evals' cold start is smaller than you think: start ugly, iterate

sarahcat21 · x · 2026-09-21

Responding to the common complaint that designing a comprehensive set of tasks and verifiers is too hard, the author argues you don't need it to get started: a few good tasks plus the means to analyze your outcomes and traces are enough to update and append over time, letting evals evolve with your agent. The quoted article's TLDR: "start ugly, write evals anyway" — understand why evals matter without burning exponential tokens. Framework cited: Agent = Model + Harness.

Original post →

More from coding & agent

coding & agent channel →