Evals are replays: store tasks in R2, sandbox in Docker, and 90% of the work is measuring
danshipper · x · 2026-10-11
Hammer (via danshipper's retweet) lays out a pragmatic approach to building agent evals without getting lost:
- Core idea: an eval is simply saving past work and replaying it against new models, prompts, or skills to see if things improved. That gives you three things: a way to pinpoint failures for your boss/users, a loop to optimize against, and a path to migrate tasks to cheaper or open-source models.
- Effort split: measuring "did it improve" is 90% of the work; infra is only 10%.
- Minimal infra: have your agent store tasks in Cloudflare R2 (cheap or free) in Harbor task format — saving the task instructions plus the folder state before the task started, so the agent sees exactly what you saw.
- Isolation: at run time the agent pulls the files and executes inside a Docker container sandbox so it can't peek at the answer.
More from coding & agent
- Anatomy: a Claude Code skill that turns any idea into an interactive machine drawing — _AustinCalvert_ · 2026-10-11
- Fine-tuned SD1.5 generates 16×16 Minecraft item sprites you can drop into resource packs — yuuki202800 · 2026-10-11
- 16x Microsoft MVP demos Jev, a PowerShell AI workflow for ranking and routing requests — dfinke · 2026-10-11
- rec-rs: a clean-room Rust rewrite of OBS for macOS using only Apple frameworks — Rasmic · 2026-10-11
- Provenance gateway turns AI agent tool calls into audit-grade, tamper-evident records — hutsonlabs · 2026-10-11
- Solo dev dilemma: self-hosted n8n or FastAPI on Cloud Run for LLM automations? — Mysterious_Profit696 · 2026-10-11