Agent evals 101: record full runs, label pass/fail, skip 1-5 scores
strickvl · x · 2026-10-08
- Record real agent runs: capture the entire session, including every LLM and tool call.
- Manually review samples: label each run pass/fail with a short critique explaining what went wrong.
- Binary labels beat 1-5 scores: a "fail" tells you something needs fixing; a "3" doesn't.
More from coding & agent
- A 4-month method for using Codex App as an information-discovery engine — jxnlco · 2026-10-08
- Handwritten text, AI-built interactions: a new golden age of technical explainers — srush_nlp · 2026-10-08
- Coinbase Institute paper: stablecoins fit agent-to-agent micropayments in AiFi era — MurrLincoln · 2026-10-08
- Depth2Depth fuses DepthAnything with RealSense for dense, metric depth maps — chrismatthieu · 2026-10-08
- 12 things engineers miss since coding died: faster models, busier teams, lost satisfaction — rseroter · 2026-10-08
- Microsoft's MAI-Code-1.1-Flash goes local in GitHub Copilot with free on-device calls — flngr · 2026-10-08