Evals are replays: store tasks in R2, sandbox in Docker, and 90% of the work is measuring

danshipper · x · 2026-10-11

Hammer (via danshipper's retweet) lays out a pragmatic approach to building agent evals without getting lost:

Original post →

More from coding & agent

coding & agent channel →